{"cells":[{"cell_type":"markdown","id":"2abd08fb","metadata":{},"source":"# Osteoarthritis is almost never written as OA\n\nThis competition ships **58 labeled studies against 4407 knee-MRI studies**. The other studies are not\nunlabeled in the usual sense: each carries a full radiology report, and the report names the findings.\nThe problem is the labels, not the pixels.\n\nThe other studies carry a full radiology report, written in several languages including Greek and Turkish,\nand the report names the findings. The exact language count is reported by the detector below.\n\nThe catch is that the report exists only in training. `test.csv` is a single column of study identifiers,\nso any deployable model has to work from images alone. That makes the report a train only teacher signal:\nif we can read the reports well enough to grade them against the 58 answer keys, we can hand any images\nonly model training targets for the 4349 report only studies.\n\nThis notebook does exactly that and nothing else. It builds a language agnostic, assertion aware rule\nlabeler from external radiology vocabulary, grades it against the 58 gold with study level bootstrap\nconfidence intervals, and emits `silver_labels.csv` for the 4349 report only studies. It does not train\nan image model and makes no leaderboard claim. Every number below is recomputed live from `train.csv` in\nthe cell that displays it.\n\nThe single sharpest result: naive keyword reading fails two independent ways at once. Findings are written\nas their consequence rather than their name, so osteoarthritis sits near AUC 0.5 until you mine osteophytes,\njoint space narrowing and chondral loss across languages. And findings are present as a word but negated,\nso a plain keyword hit counts a ruled out finding as present. Fixing both is what turns 58 answer keys into\na usable label engine."},{"cell_type":"code","execution_count":null,"id":"10932c95","metadata":{},"outputs":[],"source":"import os, re, glob, math, warnings\nimport numpy as np, pandas as pd\nimport matplotlib as mpl, matplotlib.pyplot as plt\nwarnings.filterwarnings(\"ignore\")\nmpl.rcParams[\"axes.unicode_minus\"] = False   # render signed numbers with ASCII hyphen, not U+2212\nmpl.rcParams[\"figure.dpi\"] = 120\nmpl.rcParams[\"font.size\"] = 11\nmpl.rcParams[\"axes.grid\"] = True\nmpl.rcParams[\"grid.alpha\"] = 0.25\n\ndef resolve_input():\n    cands = [\"/kaggle/input/rsna-knee-abnormality-detection\",\n             \"/kaggle/input/competitions/rsna-knee-abnormality-detection\"]\n    for c in cands:\n        if os.path.exists(os.path.join(c, \"train.csv\")):\n            return c\n    if os.path.isdir(\"/kaggle/input\"):\n        print(\"contents of /kaggle/input:\", sorted(os.listdir(\"/kaggle/input\")))\n        base = \"/kaggle/input\"; basedepth = base.rstrip(\"/\").count(\"/\")\n        for dp, dirs, files in os.walk(base):          # shallow walk, do not descend the DICOM tree\n            if \"train.csv\" in files:\n                print(\"found train.csv at\", dp); return dp\n            if dp.count(\"/\") - basedepth >= 3:\n                dirs[:] = []\n        raise FileNotFoundError(\"competition data not mounted; attach rsna-knee-abnormality-detection\")\n    return \"D:/kaggle/rsna_data\"   # local smoke only (no /kaggle/input present)\nINPUT = resolve_input()\nprint(\"INPUT =\", INPUT)\nLABELS = [\"ACL\",\"MCL\",\"Medial Meniscus\",\"Lateral Meniscus\",\"Medial OA\",\"Lateral OA\",\n          \"PF OA\",\"Effusion\",\"Synovitis\",\"Baker's\",\"Contusion\",\"Fracture\"]\nassert len(LABELS) == 12\n\ntrain = pd.read_csv(os.path.join(INPUT, \"train.csv\"))\ntrain_series = pd.read_csv(os.path.join(INPUT, \"train_series.csv\"))\ntest = pd.read_csv(os.path.join(INPUT, \"test.csv\"))\ntest_series = pd.read_csv(os.path.join(INPUT, \"test_series.csv\"))\nprint(\"train.csv       \", train.shape, \"columns:\", list(train.columns))\nprint(\"train_series.csv\", train_series.shape)\nprint(\"test.csv        \", test.shape, \"columns:\", list(test.columns))\nprint(\"test_series.csv \", test_series.shape)"},{"cell_type":"markdown","id":"60a731e1","metadata":{},"source":"## 1. The accounting: the problem is the labels, not the pixels\n\nHow much supervision actually ships, and where does the only extra signal live."},{"cell_type":"code","execution_count":null,"id":"dbfac3ed","metadata":{},"outputs":[],"source":"n_total = len(train)\nhas_report = train[\"Report\"].astype(str).str.strip().replace(\"nan\", \"\").str.len() > 0\nis_labeled = train[LABELS].notna().any(axis=1)\nn_report = int(has_report.sum())\nn_labeled = int(is_labeled.sum())\nn_lab_rep = int((is_labeled & has_report).sum())\nn_lab_norep = int((is_labeled & ~has_report).sum())\npct_labeled = 100.0 * n_labeled / n_total\n\nprint(f\"studies total                 : {n_total}\")\nprint(f\"studies with a report         : {n_report}\")\nprint(f\"studies with gold labels      : {n_labeled}   ({pct_labeled:.2f} percent of all studies)\")\nprint(f\"labeled AND has report        : {n_lab_rep}\")\nprint(f\"labeled BUT no report         : {n_lab_norep}\")\n\n# reports are train only: test has no Report column\nprint(f\"\\ntest has a Report column     : {'Report' in test.columns}\")\n\nrl = train.loc[has_report, \"Report\"].astype(str)\nclen = rl.str.len(); wlen = rl.str.split().apply(len)\nprint(f\"\\nreport length chars  min/median/mean/max : {clen.min()} / {int(clen.median())} / {clen.mean():.0f} / {clen.max()}\")\nprint(f\"report length words  min/median/mean/max : {wlen.min()} / {int(wlen.median())} / {wlen.mean():.0f} / {wlen.max()}\")"},{"cell_type":"code","execution_count":null,"id":"16cfbae8","metadata":{},"outputs":[],"source":"# FIG0  supervision gap waffle: labeled studies as a fraction of all studies\nfig, ax = plt.subplots(figsize=(9, 3.4))\ncols = 100\nrows = int(np.ceil(n_total / cols))\ngrid = np.zeros((rows, cols))\nflat = grid.ravel()\nflat[:n_labeled] = 1.0            # labeled\nflat[n_labeled:n_total] = 0.5     # report only\nflat[n_total:] = np.nan           # padding\ngrid = flat.reshape(rows, cols)\ncmap = mpl.colors.ListedColormap([\"#c7d0d8\", \"#e8a33d\"])\nax.imshow(np.where(grid == 1.0, 1, 0), cmap=cmap, aspect=\"equal\", vmin=0, vmax=1)\nax.set_xticks([]); ax.set_yticks([]); ax.grid(False)\nax.set_title(f\"Each cell is one study. Amber = gold labeled ({n_labeled} of {n_total}, {pct_labeled:.1f} percent). \"\n             f\"Grey = report only ({n_report - n_labeled}).\", fontsize=10)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"bb2ff77b","metadata":{},"source":"## 1b. What the deployable model must consume\n\nBecause the report is train only, the model that scores on the hidden test reads pixels. Each study is a\nset of DICOM series across three anatomical planes, tagged in the series metadata by plane and by whether\nthe sequence is fluid sensitive and fat suppressed. A compact montage grounds what that model sees."},{"cell_type":"code","execution_count":null,"id":"17c93197","metadata":{},"outputs":[],"source":"# series structure for train vs test\nfor name, ser in [(\"train\", train_series), (\"test\", test_series)]:\n    spg = ser.groupby(\"StudyInstanceUID\").size()\n    print(f\"{name}_series: {len(ser)} series over {ser['StudyInstanceUID'].nunique()} studies, \"\n          f\"mean {spg.mean():.1f} series/study\")\n    print(f\"   plane counts : {ser['Anatomical_Plane'].value_counts().to_dict()}\")\n    print(f\"   fluid_sensitive=1 : {int((ser['Fluid_Sensitive']==1).sum())} | \"\n          f\"fat_suppression=1 : {int((ser['Fat_Suppression']==1).sum())}\")"},{"cell_type":"code","execution_count":null,"id":"b7f2a2eb","metadata":{},"outputs":[],"source":"# FIG1  DICOM montage: one series per plane for a single study, middle slices, windowed, no PHI printed.\ndef load_series_dir(series_dir, n=3):\n    try:\n        import pydicom\n    except Exception:\n        return []\n    files = sorted(glob.glob(os.path.join(series_dir, \"*.dcm\")))\n    if not files:\n        return []\n    # sort by InstanceNumber when present\n    def inst(f):\n        try:\n            return int(pydicom.dcmread(f, stop_before_pixels=True).get(\"InstanceNumber\", 0))\n        except Exception:\n            return 0\n    files = sorted(files, key=inst)\n    mid = len(files) // 2\n    pick = files[max(0, mid - n // 2): max(0, mid - n // 2) + n]\n    imgs = []\n    for f in pick:\n        try:\n            ds = pydicom.dcmread(f)\n            a = ds.pixel_array.astype(np.float32)\n            if str(ds.get(\"PhotometricInterpretation\", \"\")) == \"MONOCHROME1\":\n                a = a.max() - a\n            lo, hi = np.percentile(a, [1, 99])\n            a = np.clip((a - lo) / max(hi - lo, 1e-6), 0, 1)\n            imgs.append(a)\n        except Exception:\n            pass\n    return imgs\n\nmontage_done = False\ntry:\n    planes = [\"Sagittal\", \"Axial\", \"Coronal\"]\n    # choose a study that has all three planes\n    ser = train_series if os.path.isdir(os.path.join(INPUT, \"train_series\")) else test_series\n    base = \"train_series\" if ser is train_series else \"test_series\"\n    piv = ser.groupby(\"StudyInstanceUID\")[\"Anatomical_Plane\"].apply(lambda s: set(s))\n    cand = [sid for sid, ps in piv.items() if set(planes).issubset(ps)]\n    if cand:\n        sid = cand[0]\n        rows_imgs = []\n        for pl in planes:\n            row = ser[(ser.StudyInstanceUID == sid) & (ser.Anatomical_Plane == pl)]\n            if len(row) == 0:\n                rows_imgs.append([]); continue\n            series_id = row.iloc[0][\"SeriesInstanceUID\"]\n            rows_imgs.append(load_series_dir(os.path.join(INPUT, base, sid, series_id), n=3))\n        ncol = max((len(r) for r in rows_imgs), default=0)\n        if ncol > 0:\n            fig, axes = plt.subplots(3, ncol, figsize=(2.4 * ncol, 7.2))\n            axes = np.atleast_2d(axes)\n            for i, (pl, imgs) in enumerate(zip(planes, rows_imgs)):\n                for j in range(ncol):\n                    ax = axes[i, j]; ax.axis(\"off\"); ax.grid(False)\n                    if j < len(imgs):\n                        ax.imshow(imgs[j], cmap=\"gray\")\n                    if j == 0:\n                        ax.set_title(pl, loc=\"left\", fontsize=11)\n            fig.suptitle(\"One knee study across three planes (windowed, identifiers not shown)\", y=0.99)\n            plt.tight_layout(); plt.show()\n            montage_done = True\nexcept Exception as e:\n    print(\"montage skipped:\", repr(e))\nif not montage_done:\n    print(\"DICOM montage unavailable in this environment (full image data is present on Kaggle).\")"},{"cell_type":"markdown","id":"31463fc4","metadata":{},"source":"## 1c. Reading the pixels: windowing, sequence contrast, and a volume sweep\n\nThree checks on the raw imaging before any labels: the exact windowing the pipeline uses, that the sequence\nmetadata is visible in the pixels, and how one series sweeps through the joint. Pure pixels, no pathology\nclaim on any single slice."},{"cell_type":"code","execution_count":null,"id":"ce68d3c0","metadata":{},"outputs":[],"source":"import pydicom\nBASE_TRAIN = \"train_series\" if os.path.isdir(os.path.join(INPUT, \"train_series\")) else None\ndef window_slice(ds):\n    a = ds.pixel_array.astype(np.float32)\n    if str(ds.get(\"PhotometricInterpretation\", \"\")) == \"MONOCHROME1\": a = a.max() - a\n    lo, hi = np.percentile(a, [1, 99]); return np.clip((a - lo) / max(hi - lo, 1e-6), 0, 1)\ndef resize2d(a, s=300):\n    yi = np.clip(np.round(np.linspace(0, a.shape[0]-1, s)).astype(int), 0, a.shape[0]-1)\n    xi = np.clip(np.round(np.linspace(0, a.shape[1]-1, s)).astype(int), 0, a.shape[1]-1)\n    return a[yi][:, xi]\ndef sdir(base, sid, serid): return os.path.join(INPUT, base, sid, serid)\ndef list_dcm(d): return sorted(glob.glob(os.path.join(d, \"*.dcm\")))\ndef by_instance(files):\n    def k(f):\n        try: return int(pydicom.dcmread(f, stop_before_pixels=True).get(\"InstanceNumber\", 0))\n        except Exception: return 0\n    return sorted(files, key=k)\ndef central_windowed(d):\n    files = list_dcm(d)\n    if not files: return None\n    files = by_instance(files)\n    try: return window_slice(pydicom.dcmread(files[len(files)//2]))\n    except Exception: return None\ndef pick_series_for(rows, plane, need_fs=False, need_fsat=False):\n    sub = rows if plane == \"any\" else rows[rows.Anatomical_Plane == plane]\n    if need_fs: sub = sub[sub.Fluid_Sensitive == 1]\n    if need_fsat: sub = sub[sub.Fat_Suppression == 1]\n    if len(sub) == 0: return None\n    sub = sub.sort_values([\"Fluid_Sensitive\", \"Fat_Suppression\"], ascending=False)\n    return sub.iloc[0][\"SeriesInstanceUID\"]\nis_labeled = train[LABELS].notna().any(axis=1)\ngold = train[is_labeled].copy().reset_index(drop=True)\nprint(\"imaging helpers ready | BASE_TRAIN:\", BASE_TRAIN, \"| gold studies:\", len(gold))"},{"cell_type":"code","execution_count":null,"id":"f1d2c1f6","metadata":{},"outputs":[],"source":"# NEW-A  photometric and windowing forensics (raw uint16 -> percentile window -> histogram). No pathology claim.\ntry:\n    done = False\n    if BASE_TRAIN:\n        for _, r in train_series.head(60).iterrows():\n            files = list_dcm(sdir(BASE_TRAIN, r.StudyInstanceUID, r.SeriesInstanceUID))\n            if not files: continue\n            ds = pydicom.dcmread(files[len(files)//2])\n            if str(ds.get(\"PhotometricInterpretation\", \"\")) == \"MONOCHROME1\": continue\n            raw = ds.pixel_array.astype(np.float32)\n            lo, hi = np.percentile(raw, [1, 99])\n            fig, ax = plt.subplots(1, 3, figsize=(14, 4.5))\n            ax[0].imshow(resize2d(raw, 320), cmap=\"gray\"); ax[0].set_axis_off()\n            ax[0].set_title(f\"raw uint16  min {int(raw.min())}  max {int(raw.max())}\")\n            ax[1].imshow(resize2d(window_slice(ds), 320), cmap=\"gray\", vmin=0, vmax=1); ax[1].set_axis_off()\n            ax[1].set_title(\"windowed 1 to 99 pct, scaled 0 to 1\")\n            ax[2].hist(raw.ravel(), bins=200, log=True, color=\"#3d7ea8\")\n            ax[2].axvline(lo, color=\"#c0603a\"); ax[2].axvline(hi, color=\"#c0603a\"); ax[2].grid(False)\n            ax[2].set_title(f\"intensity histogram (1 pct = {int(lo)}, 99 pct = {int(hi)})\")\n            ax[2].set_xlabel(\"raw intensity\"); ax[2].set_ylabel(\"log count\")\n            plt.tight_layout(); plt.show()\n            print(\"Display normalization only; different windowing changes apparent conspicuity.\")\n            done = True; break\n    if not done: print(\"NEW-A: DICOM not available in this environment (renders on Kaggle)\")\nexcept Exception as e:\n    print(\"NEW-A skipped:\", repr(e))"},{"cell_type":"code","execution_count":null,"id":"34a425c6","metadata":{},"outputs":[],"source":"# NEW-B  sequence contrast on one knee: same plane, central slice across Fluid_Sensitive x Fat_Suppression.\ntry:\n    done = False\n    if BASE_TRAIN and len(gold):\n        cand = (gold[gold.get(\"Effusion\", 0) == 1][\"StudyInstanceUID\"].tolist()\n                if \"Effusion\" in gold.columns else []) + gold[\"StudyInstanceUID\"].tolist()\n        best = None\n        for sid in cand[:40]:\n            rws = train_series[train_series.StudyInstanceUID == sid]\n            for pl in [\"Axial\", \"Sagittal\", \"Coronal\"]:\n                sub = rws[rws.Anatomical_Plane == pl]\n                if len(set(zip(sub.Fluid_Sensitive, sub.Fat_Suppression))) >= 2:\n                    best = (sid, pl); break\n            if best: break\n        if best:\n            sid, pl = best\n            rws = train_series[(train_series.StudyInstanceUID == sid) & (train_series.Anatomical_Plane == pl)]\n            fig, axes = plt.subplots(2, 2, figsize=(8, 8))\n            for fi, fs in enumerate([1, 0]):\n                for gi, fsat in enumerate([1, 0]):\n                    ax = axes[fi, gi]; ax.set_axis_off()\n                    ss = rws[(rws.Fluid_Sensitive == fs) & (rws.Fat_Suppression == fsat)]\n                    a = central_windowed(sdir(BASE_TRAIN, sid, ss.iloc[0].SeriesInstanceUID)) if len(ss) else None\n                    if a is None:\n                        ax.text(0.5, 0.5, \"not acquired\", ha=\"center\", va=\"center\"); continue\n                    ax.imshow(resize2d(a, 300), cmap=\"gray\", vmin=0, vmax=1)\n                    ax.set_title(f\"FS {fs}  FatSat {fsat}\")\n                    ax.text(4, 20, f\"mean {a.mean():.2f}  bright {(a > 0.8).mean():.2f}\", color=\"yellow\", fontsize=8)\n            fig.suptitle(f\"Same knee, {pl} plane, sequence contrast (central slice, shared scale, not co-registered)\")\n            plt.tight_layout(); plt.show()\n            print(\"Bright fraction is a whole-slice statistic, not an effusion segmentation and not a detector; not verified at pixel level.\")\n            done = True\n    if not done: print(\"NEW-B: no study with two sequence combos available here (renders on Kaggle)\")\nexcept Exception as e:\n    print(\"NEW-B skipped:\", repr(e))"},{"cell_type":"code","execution_count":null,"id":"6ecea850","metadata":{},"outputs":[],"source":"# NEW-D  volume sweep: ten evenly spaced slices through one canonical series, ordered by InstanceNumber.\ntry:\n    done = False\n    if BASE_TRAIN and len(gold):\n        sid = gold[\"StudyInstanceUID\"].iloc[0]\n        rws = train_series[train_series.StudyInstanceUID == sid]\n        serid = pick_series_for(rws, \"Sagittal\", need_fs=True) or pick_series_for(rws, \"Sagittal\")                 or (rws.iloc[0].SeriesInstanceUID if len(rws) else None)\n        files = by_instance(list_dcm(sdir(BASE_TRAIN, sid, serid))) if serid else []\n        if files:\n            M = min(10, len(files)); idx = np.unique(np.round(np.linspace(0, len(files)-1, M)).astype(int))\n            fig, axes = plt.subplots(2, 5, figsize=(15, 6)); axes = axes.ravel()\n            for k in range(len(axes)): axes[k].set_axis_off()\n            for k, ii in enumerate(idx):\n                try:\n                    a = window_slice(pydicom.dcmread(files[int(ii)]))\n                    axes[k].imshow(resize2d(a, 300), cmap=\"gray\", vmin=0, vmax=1)\n                    axes[k].set_title(f\"slice {int(ii)+1} of {len(files)}\", fontsize=9)\n                except Exception: pass\n            fig.suptitle(\"Sagittal sweep (index order; anatomical direction not decoded)\")\n            plt.tight_layout(); plt.show()\n            done = True\n    if not done: print(\"NEW-D: DICOM not available in this environment (renders on Kaggle)\")\nexcept Exception as e:\n    print(\"NEW-D skipped:\", repr(e))"},{"cell_type":"markdown","id":"9b1321b1","metadata":{},"source":"## 2. The target: multiple languages and a multi label answer key\n\nA crude offline language detector, run only to show scale. The undetermined bucket is large and includes\nscripts the simple heuristic does not cover, so the count is a lower bound."},{"cell_type":"code","execution_count":null,"id":"435e430a","metadata":{},"outputs":[],"source":"# crude language guess by stopword hits (offline, heuristic; a lower bound on language count)\nLANG = {\n \"English\":[\" the \",\" and \",\" is \",\" no \",\" tear \",\" medial \",\" lateral \",\" knee \",\" intact \",\" normal \"],\n \"Spanish\":[\" el \",\" la \",\" de \",\" rotura \",\" menisco \",\" rodilla \",\" derrame \",\" sin \",\" con \",\" ligamento \"],\n \"German\":[\" und \",\" mit \",\" ohne \",\" der \",\" die \",\" kein \",\" kniegelenk \",\" riss \",\" unauffallig \"],\n \"Dutch\":[\" en \",\" met \",\" zonder \",\" van \",\" rechts \",\" meniscus \",\" knie \",\" geen \",\" ongestoord \"],\n \"French\":[\" le \",\" et \",\" avec \",\" sans \",\" genou \",\" menisque \",\" rupture \",\" pas de \"],\n \"Italian\":[\" il \",\" con \",\" senza \",\" del \",\" ginocchio \",\" rottura \",\" non \"],\n \"Portuguese\":[\" do \",\" com \",\" sem \",\" joelho \",\" rotura \",\" nao \"],\n \"Turkish\":[\" ve \",\" ile \",\" yok \",\" diz \",\" menisk \",\" normal \",\" bulgular \"],\n \"Greek\":[\"\\u03b7 \",\"\\u03bc\\u03b5 \",\"\\u03c4\\u03bf\\u03c5 \",\"\\u03b3\\u03bf\\u03bd\",\"\\u03bc\\u03b7\\u03bd\\u03b9\\u03c3\\u03ba\"],\n}\ndef guess_lang(t):\n    s = \" \" + str(t).lower() + \" \"\n    best, bs = \"Undetermined\", 1\n    for k, ws in LANG.items():\n        c = sum(s.count(w) for w in ws)\n        if c > bs: best, bs = k, c\n    return best\nlang = train.loc[has_report, \"Report\"].apply(guess_lang)\nlc = lang.value_counts()\nn_detected = int((lc.index != \"Undetermined\").sum())\nprint(f\"languages with a positive detection: {n_detected}\")\nfor k, v in lc.items():\n    print(f\"  {k:14s} {v:5d}  ({100.0*v/len(lang):.1f} percent)\")"},{"cell_type":"code","execution_count":null,"id":"8d900975","metadata":{},"outputs":[],"source":"# FIG2  language distribution (crude detector)\nfig, ax = plt.subplots(figsize=(8, 4.2))\norder = lc.sort_values()\ncolors = [\"#9aa7b1\" if k == \"Undetermined\" else \"#3d7ea8\" for k in order.index]\nax.barh(range(len(order)), order.values, color=colors)\nax.set_yticks(range(len(order))); ax.set_yticklabels(order.index)\nax.set_xlabel(\"studies\"); ax.set_title(f\"Report languages, crude detector, at least {n_detected} detected\")\nfor i, v in enumerate(order.values):\n    ax.text(v + max(order.values)*0.01, i, str(int(v)), va=\"center\", fontsize=9)\nplt.tight_layout(); plt.show()"},{"cell_type":"code","execution_count":null,"id":"cf2488dd","metadata":{},"outputs":[],"source":"# gold label frequency and 12x12 co-occurrence on the 58 gold\ngold = train[is_labeled].copy().reset_index(drop=True)\nG = gold[LABELS].fillna(0).astype(int)\nnpos = G.sum(axis=0)\nmean_pos = G.sum(axis=1).mean()\nprint(f\"gold studies: {len(G)}   mean positives per study: {mean_pos:.2f}\")\nprint(\"per label positive count (of 58):\")\nfor lab in LABELS:\n    print(f\"  {lab:18s} {int(npos[lab]):3d}\")\nco = np.zeros((12, 12), int)\nfor i in range(12):\n    for j in range(12):\n        co[i, j] = int(((G[LABELS[i]] == 1) & (G[LABELS[j]] == 1)).sum())\npairs = sorted(((co[i, j], LABELS[i], LABELS[j]) for i in range(12) for j in range(i+1, 12)), reverse=True)\nprint(\"\\ntop co-occurrence pairs:\")\nfor c, a, b in pairs[:6]:\n    print(f\"  {c:2d}  {a} + {b}\")"},{"cell_type":"code","execution_count":null,"id":"616342e1","metadata":{},"outputs":[],"source":"# FIG3  gold label frequency bar + co-occurrence heatmap\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(13.5, 5.2), gridspec_kw={\"width_ratios\":[1, 1.15]})\noi = np.argsort(npos.values)\nax1.barh(range(12), npos.values[oi], color=\"#c0603a\")\nax1.set_yticks(range(12)); ax1.set_yticklabels([LABELS[i] for i in oi])\nax1.set_xlabel(\"positive studies (of 58)\")\nax1.set_title(f\"Gold label frequency (mean {mean_pos:.1f} positives per study)\")\nfor i, v in enumerate(npos.values[oi]):\n    ax1.text(v + 0.3, i, str(int(v)), va=\"center\", fontsize=9)\nim = ax2.imshow(co, cmap=\"YlOrRd\")\nax2.set_xticks(range(12)); ax2.set_xticklabels(LABELS, rotation=90, fontsize=8)\nax2.set_yticks(range(12)); ax2.set_yticklabels(LABELS, fontsize=8)\nax2.grid(False)\neff = LABELS.index(\"Effusion\")\nax2.add_patch(plt.Rectangle((-0.5, eff-0.5), 12, 1, fill=False, edgecolor=\"#1f4e79\", lw=2))\nax2.set_title(\"Co-occurrence counts, n=58 (Effusion row outlined)\")\nfor i in range(12):\n    for j in range(12):\n        if co[i, j] > 0:\n            ax2.text(j, i, co[i, j], ha=\"center\", va=\"center\", fontsize=6,\n                     color=\"white\" if co[i, j] > co.max()*0.6 else \"black\")\nfig.colorbar(im, ax=ax2, fraction=0.046)\nplt.tight_layout(); plt.show()\nprint(\"Caveat: n=58 is tiny; co-occurrence is descriptive of this enriched sample, not a test time prior.\")"},{"cell_type":"markdown","id":"aa8fc636","metadata":{},"source":"## 3. The labeler\n\nA language agnostic, assertion aware rule labeler. It fires a 12 concept lexicon built from external\nmultilingual radiology vocabulary on every scope of the report, decides for each mention whether it is\naffirmed, negated or uncertain from the surrounding cue words in six plus languages, and aggregates a soft\nscore per label. It uses no machine translation and no report text at inference. The full module is below,\nso the notebook is self contained and forkable."},{"cell_type":"code","execution_count":null,"id":"c2744100","metadata":{},"outputs":[],"source":"# RSNA Knee report labeler: language-agnostic, assertion-aware (ConText), no machine translation.\n# Lexicon built from EXTERNAL multilingual radiology terminology (standard finding names), not by\n# fitting to which tokens appear in the 58 gold. Emits soft per-label scores + abstain flags.\n# Pure stdlib + numpy + pandas. Importable; used identically by the notebook and by local validation.\nimport re, unicodedata\nimport numpy as np\n\nLABELS = [\"ACL\",\"MCL\",\"Medial Meniscus\",\"Lateral Meniscus\",\"Medial OA\",\"Lateral OA\",\n          \"PF OA\",\"Effusion\",\"Synovitis\",\"Baker's\",\"Contusion\",\"Fracture\"]\n\ndef norm(t):\n    t = str(t).lower()\n    t = unicodedata.normalize(\"NFKD\", t)\n    t = \"\".join(c for c in t if not unicodedata.combining(c))\n    return re.sub(r\"\\s+\", \" \", t)\n\n# ---- assertion cues (multilingual), matched inside a bounded scope ----\nNEG_PRE = [  # negation appearing BEFORE the finding\n    \"no \",\"not \",\"non \",\"without\",\"w/o\",\"absence of\",\"absent\",\"negative for\",\"no evidence\",\n    \"no sign\",\"free of\",\"rule out\",\"ruled out\",\"r/o\",\n    \"sin \",\"ausencia\",\"ausente\",\"descarta\",\"no hay\",\"no se observ\",\"no se identific\",\"sin signos\",\n    \"sin evidencia\",\"no existe\",\"no presenta\",\n    \"kein\",\"keine\",\"ohne\",\"nicht\",\"negativ\",\n    \"geen\",\"zonder\",\"niet\",\"negatief\",\n    \"pas de\",\"pas d\",\"aucun\",\"absence\",\n    \"senza\",\"assenza\",\"nessun\",\"non \",\n    \"sem \",\"nao \",\"yok\",\"olmayan\",\"izlenmedi\",\"saptanmadi\",\"gozlenmedi\",\n]\nNORMAL_POST = [  # normality / intactness AFTER the structure implies negative\n    \"intact\",\"normal\",\"unremarkable\",\"preserved\",\"conserved\",\"within normal\",\n    \"integr\",\"conservad\",\"normale\",\"normales\",\"sin alteracion\",\"sin particularidad\",\n    \"unauffallig\",\"regelrecht\",\"erhalten\",\"intakt\",\n    \"normaal\",\"ongestoord\",\n    \"sans particularite\",\"integre\",\n    \"normaldir\",\"dogal\",\"olagan\",\n]\nUNC = [  # uncertainty / hedge\n    \"possible\",\"possibly\",\"cannot exclude\",\"can not exclude\",\"suspicious\",\"suspected\",\"probable\",\n    \"likely\",\"questionable\",\"equivocal\",\"may represent\",\"concerning for\",\n    \"posible\",\"no se puede excluir\",\"sospech\",\"probable\",\"dudos\",\"sugestiv\",\"sugiere\",\n    \"moglich\",\"nicht auszuschliessen\",\"verdacht\",\"v.a\",\"dd \",\n    \"mogelijk\",\"verdenking\",\"waarschijnlijk\",\n    \"possible\",\"probable\",\"suspect\",\n    \"supheli\",\"olabilir\",\"suphesi\",\n]\nADVERSATIVE = [\"but \",\"however\",\"although\",\"though\",\"pero \",\"aunque\",\"sin embargo\",\"jedoch\",\"aber \",\n               \"echter\",\"maar \",\"mais \",\"cependant\",\"ma \",\"tuttavia\",\"porem\",\"mas \"]\n\n# ---- indication / clinical-question headers to DROP (state suspicion, not findings) ----\nINDIC_HDR = [\"indication\",\"clinical\",\"history\",\"reason\",\"question\",\"antecedent\",\"motivo\",\"clinica\",\n             \"vraagstelling\",\"inlichting\",\"fragestellung\",\"klinische\",\"anamnes\",\"indicac\",\"hikaye\",\n             \"klinik\",\"diagnostische vraag\",\"clinical inlichtingen\"]\nIMPRESS_HDR = [\"impression\",\"conclusion\",\"assessment\",\"opinion\",\"impresion\",\"conclusion\",\"besluit\",\n               \"beurteilung\",\"zusammenfassung\",\"conclusie\",\"sonuc\",\"diagnostico\",\"impressie\",\"fazit\"]\n\ndef split_sentences(text):\n    return [s.strip() for s in re.split(r\"[.\\n;:]+\", text) if s.strip()]\n\ndef split_scopes(clause):\n    # terminate scope at adversative conjunctions so \"no effusion but meniscus tear\" scopes correctly\n    parts = [clause]\n    for adv in ADVERSATIVE:\n        out = []\n        for p in parts:\n            out.extend(p.split(adv))\n        parts = out\n    return [p.strip() for p in parts if p.strip()]\n\ndef assertion(scope, term):\n    \"\"\"Return AFFIRMED / NEGATED / UNCERTAIN for a term mention inside its scope.\"\"\"\n    i = scope.find(term)\n    if i < 0:\n        return None\n    before = scope[max(0, i-55):i]\n    after = scope[i+len(term): i+len(term)+40]\n    if any(u in scope for u in UNC):\n        # uncertainty only if the hedge is near the term (same scope already bounds it)\n        if any(u in before or u in after for u in UNC):\n            return \"UNCERTAIN\"\n    if any(n in before for n in NEG_PRE):\n        return \"NEGATED\"\n    if any(x in after for x in NORMAL_POST):\n        return \"NEGATED\"\n    return \"AFFIRMED\"\n\n# ---- concept lexicon (external radiology vocabulary) ----\n# TEAR words for anatomy-based classes (ACL/MCL/menisci)\nTEAR = [norm(x) for x in [\"tear\",\"torn\",\"rupture\",\"ruptured\",\"rupt\",\"disrupt\",\"disruption\",\"rotura\",\n    \"roto\",\"rota\",\"desgarro\",\"ruptura\",\"desinsercion\",\"riss\",\"gerissen\",\"ruptur\",\"scheur\",\"gescheurd\",\n    \"ruptuur\",\"dechirure\",\"dechir\",\"rottura\",\"lacerazione\",\"yirtik\",\"yirtig\",\"rexis\",\"ρηξη\"]]\n# meniscus signals that should DEMOTE from a true tear\nMENISC_DEMOTE = [norm(x) for x in [\"grade 1\",\"grade i\",\"grade 2\",\"grade ii\",\"grado 1\",\"grado 2\",\n    \"intrasubstance\",\"intrasustancia\",\"mucoid\",\"mixoide\",\"degenerac\",\"degenerative signal\",\n    \"no surface\",\"does not contact\",\"sin contactar\",\"without surface\",\"gebni\"]]\n\ndef clause_has(scope, terms):\n    return any(t in scope for t in terms)\n\n# structure regex fragments per compartment/ligament (accent-folded)\nS_ACL = [norm(x) for x in [\"acl\",\"anterior cruciate\",\"cruzado anterior\",\"lca\",\"vorderes kreuzband\",\"vkb\",\n    \"voorste kruisband\",\"ligament croise anterieur\",\"crociato anteriore\",\"on capraz bag\",\"kreuzband vorder\"]]\nS_MCL = [norm(x) for x in [\"mcl\",\"medial collateral\",\"colateral medial\",\"colateral interno\",\"lli\",\"lcm\",\n    \"ligamento lateral interno\",\"innenband\",\"mediales seitenband\",\"mediale collaterale\",\"medial kollateral\",\n    \"ligament collateral medial\",\"ligament lateral interne\",\"ic bag\",\"medial band\"]]\nS_MMEN = [norm(x) for x in [\"medial meniscus\",\"menisco interno\",\"menisco medial\",\"innenmeniskus\",\n    \"medialer meniskus\",\"mediale meniscus\",\"binnenmeniscus\",\"menisque interne\",\"menisque medial\",\n    \"menisco mediale\",\"medial meniskus\",\"ic menisk\"]]\nS_LMEN = [norm(x) for x in [\"lateral meniscus\",\"menisco externo\",\"menisco lateral\",\"aussenmeniskus\",\n    \"lateraler meniskus\",\"laterale meniscus\",\"buitenmeniscus\",\"menisque externe\",\"menisque lateral\",\n    \"menisco laterale\",\"lateral meniskus\",\"dis menisk\"]]\n\n# finding-based single concepts\nC_EFFUSION = [norm(x) for x in [\"effusion\",\"joint effusion\",\"joint fluid\",\"synovial fluid\",\"derrame\",\n    \"derrame articular\",\"erguss\",\"gelenkerguss\",\"epanchement\",\"versamento\",\"effusie\",\"gewrichtsvocht\",\n    \"hydrops\",\"hidrops\",\"intraarticular fluid\",\"intra-articular fluid\",\"liquido articular\",\n    \"liquido intraarticular\",\"liquido intra-articular\",\"eklem sivisi\",\"hemartros\"]]\nC_SYNOV = [norm(x) for x in [\"synovitis\",\"synovial thickening\",\"synovial hypertroph\",\"sinovitis\",\n    \"synovialitis\",\"synovite\",\"sinovite\",\"synoviale verdikking\",\"engrosamiento sinovial\",\n    \"proliferacion sinovial\",\"hypertrophie synoviale\",\"hoffitis\",\"sinovial hipertrofia\"]]\nC_BAKER = [norm(x) for x in [\"baker\",\"popliteal cyst\",\"quiste de baker\",\"quiste poplite\",\"quiste popliteo\",\n    \"bakerzyste\",\"poplitealzyste\",\"bakercyste\",\"poplitea cyste\",\"kyste de baker\",\"kyste poplite\",\n    \"cisti di baker\",\"cisti poplitea\",\"baker kisti\",\"popliteal kist\"]]\nC_CONTUS = [norm(x) for x in [\"contusion\",\"bone bruise\",\"bone marrow edema\",\"bone marrow oedema\",\n    \"marrow edema\",\"marrow oedema\",\"contusion osea\",\"edema oseo\",\"edema medular\",\"edema de medula\",\n    \"knochenmarkodem\",\"botcontusie\",\"botoedeem\",\"beenmergoedeem\",\"contusion osseuse\",\"oedeme osseux\",\n    \"oedeme medullaire\",\"subchondral edema\",\"subchondral bone marrow edema\",\"bone edema\",\"kemik odem\",\n    \"kontuzyon\",\"edema oseo subcondral\"]]\nC_FRACT = [norm(x) for x in [\"fracture\",\"fractura\",\"fraktur\",\"fractuur\",\"frattura\",\"fratura\",\"avulsion\",\n    \"insufficiency fracture\",\"stress fracture\",\"subchondral fracture\",\"fisura osea\",\"kirik\",\"fissur oss\"]]\n\n# OA consequence vocabulary (compartment assigned separately)\nOA_ANY = [norm(x) for x in [\"osteoarthr\",\"artrosis\",\"artrose\",\"gonartros\",\"gonarthros\",\"arthrose\",\n    \"osteophyt\",\"osteofit\",\"osteofyt\",\"joint space narrow\",\"joint space loss\",\"jsn\",\"joint-space narrow\",\n    \"pinzamiento\",\"estrechamiento\",\"reduccion de la interlinea\",\"interlinea articular reducida\",\n    \"chondral loss\",\"cartilage loss\",\"cartilage thinning\",\"chondrosis\",\"condrosis\",\"condromalacia\",\n    \"condropatia\",\"chondromalacia\",\"chondropath\",\"ulcera condral\",\"ulceras condral\",\"condral defect\",\n    \"chondral defect\",\"degenerativ\",\"degeneratief\",\"mucoid degeneration\",\"subchondral cyst\",\n    \"subchondral sclerosis\",\"esclerosis subcondral\",\"kellgren\",\"degenerative change\",\"kraakbeenverlies\"]]\nOA_TRICOMP = [norm(x) for x in [\"tricompartmental\",\"three compartment\",\"three-compartment\",\"pancompartmental\",\n    \"pangonarthros\",\"incipient oa of all\",\"oa of all three\",\"tricompartimental\",\"all three compartment\"]]\nCTX_MED = [norm(x) for x in [\"medial\",\"interno\",\"interna\",\"medialis\",\"binnen\",\"femorotibial medial\",\n    \"femorotibial interna\",\"medial compartment\",\"compartimento medial\",\"mediale femorotibiaal\",\"mediaal\"]]\nCTX_LAT = [norm(x) for x in [\"lateral\",\"externo\",\"externa\",\"buiten\",\"femorotibial lateral\",\n    \"femorotibial externa\",\"lateral compartment\",\"compartimento lateral\",\"laterale femorotibiaal\"]]\nCTX_PF = [norm(x) for x in [\"patellofemoral\",\"femoropatel\",\"retropatel\",\"patelofemoral\",\"rotulian\",\n    \"patellar\",\"patela \",\"rotula\",\"troclea\",\"trochlea\",\"femoro-patellaire\",\"femoropatellar\",\"patello-femoral\"]]\nOA_GENERIC = [norm(x) for x in [\"gonartros\",\"gonarthros\",\"osteoarthr\",\"artrosis\",\"artrose\",\"arthrose\"]]\n\ndef score_report(report, use_negation=True, use_oa_vocab=True):\n    \"\"\"Return dict label -> evidence score e (real). Higher = more likely positive.\n    use_negation/use_oa_vocab toggles support the ablation ladder.\"\"\"\n    raw = norm(report)\n    lines = [l.strip() for l in re.split(r\"[\\n]+\", raw) if l.strip()]\n    # section weighting: mark impression lines x1.5, drop indication lines\n    weighted = []  # (text, weight)\n    in_indic = False\n    for l in lines:\n        head = l[:40]\n        if any(h in head for h in INDIC_HDR) and (\":\" in l[:40] or len(l) < 60):\n            in_indic = True\n            continue\n        if any(h in head for h in IMPRESS_HDR):\n            in_indic = False\n            weighted.append((l, 1.5)); continue\n        # a findings/results header ends indication\n        if any(h in head for h in [\"finding\",\"result\",\"hallazgo\",\"bevinding\",\"befund\",\"bulgular\",\"resultat\"]):\n            in_indic = False; weighted.append((l, 1.0)); continue\n        if in_indic:\n            continue\n        weighted.append((l, 1.0))\n    if not weighted:\n        weighted = [(raw, 1.0)]\n\n    acc = {l: 0.0 for l in LABELS}\n    W_AFF, W_UNC, W_NEG = 1.0, 0.4, 1.0\n\n    def add(label, state, w):\n        if state is None:\n            return\n        if not use_negation:\n            # crude keyword-presence baseline: any detected mention counts as positive\n            acc[label] += W_AFF * w\n            return\n        if state == \"AFFIRMED\": acc[label] += W_AFF * w\n        elif state == \"UNCERTAIN\": acc[label] += W_UNC * w\n        elif state == \"NEGATED\": acc[label] -= W_NEG * w\n\n    for text, w in weighted:\n        for clause in split_sentences(text):\n            for scope in split_scopes(clause):\n                # anatomy-based classes: structure + tear in scope\n                for label, struct in ((\"ACL\",S_ACL),(\"MCL\",S_MCL),\n                                      (\"Medial Meniscus\",S_MMEN),(\"Lateral Meniscus\",S_LMEN)):\n                    if clause_has(scope, struct) and clause_has(scope, TEAR):\n                        # demote meniscus grade-1/2/intrasubstance/mucoid to uncertain\n                        st = None\n                        for s in struct:\n                            if s in scope:\n                                st = assertion(scope, s); break\n                        if st is None:\n                            continue\n                        if label in (\"Medial Meniscus\",\"Lateral Meniscus\") and st == \"AFFIRMED\" \\\n                           and clause_has(scope, MENISC_DEMOTE):\n                            st = \"UNCERTAIN\"\n                        add(label, st, w)\n                # finding-based single concepts\n                for label, terms in ((\"Effusion\",C_EFFUSION),(\"Synovitis\",C_SYNOV),\n                                     (\"Baker's\",C_BAKER),(\"Contusion\",C_CONTUS),(\"Fracture\",C_FRACT)):\n                    for t in terms:\n                        if t in scope:\n                            add(label, assertion(scope, t), w); break\n                # OA classes\n                if use_oa_vocab:\n                    if clause_has(scope, OA_TRICOMP):\n                        for lab in (\"Medial OA\",\"Lateral OA\",\"PF OA\"):\n                            acc[lab] += W_AFF * w\n                    for o in OA_ANY:\n                        if o in scope:\n                            st = assertion(scope, o)\n                            hit_ctx = False\n                            if clause_has(scope, CTX_MED): add(\"Medial OA\", st, w); hit_ctx=True\n                            if clause_has(scope, CTX_LAT): add(\"Lateral OA\", st, w); hit_ctx=True\n                            if clause_has(scope, CTX_PF):  add(\"PF OA\", st, w); hit_ctx=True\n                            if not hit_ctx and any(g in scope for g in OA_GENERIC):\n                                # bare knee OA -> low-weight signal to all three\n                                for lab in (\"Medial OA\",\"Lateral OA\",\"PF OA\"):\n                                    add(lab, st, 0.5*w)\n                            break\n                else:\n                    # naive OA: only the literal words \"osteoarthritis/oa/artrosis\" with compartment\n                    for o in [norm(x) for x in [\"osteoarthr\",\"artrosis\",\"oa \"]]:\n                        if o in scope:\n                            st = assertion(scope, o)\n                            if clause_has(scope, CTX_MED): add(\"Medial OA\", st, w)\n                            if clause_has(scope, CTX_LAT): add(\"Lateral OA\", st, w)\n                            if clause_has(scope, CTX_PF):  add(\"PF OA\", st, w)\n                            break\n    return acc\n\ndef sigmoid(x): return 1.0 / (1.0 + np.exp(-x))\n\ndef evidence_frame(reports, use_negation=True, use_oa_vocab=True):\n    \"\"\"Return a DataFrame of raw evidence scores (columns=LABELS) for a series of reports.\"\"\"\n    import pandas as pd\n    rows = [score_report(r, use_negation, use_oa_vocab) for r in reports]\n    return pd.DataFrame(rows)[LABELS]\n"},{"cell_type":"markdown","id":"aa32e7b5","metadata":{},"source":"### 3a. Failure mode one: the finding is written as its consequence\n\nRadiologists rarely write the word osteoarthritis. They write osteophytes, joint space narrowing, chondral\nloss, chondrosis, gonarthrose, mucoid degeneration, and they name the compartment. A labeler that only\nlooks for the literal disease name is near chance on the osteoarthritis labels. Adding the consequence\nvocabulary and the rule that tricompartmental fires all three compartments is the single largest lever on\nthe macro score, because the metric is the unweighted mean of twelve AUCs and this resurrects dead labels."},{"cell_type":"code","execution_count":null,"id":"9a4ffc23","metadata":{},"outputs":[],"source":"from sklearn.metrics import roc_auc_score, matthews_corrcoef, cohen_kappa_score\ny = gold[LABELS].fillna(0).astype(int).values\nreports = gold[\"Report\"].tolist()\n\ndef macro_auc(Y, S):\n    a = []\n    for j in range(Y.shape[1]):\n        yj = Y[:, j]\n        a.append(roc_auc_score(yj, S[:, j]) if len(set(yj)) > 1 else np.nan)\n    return np.nanmean(a), np.array(a)\n\nS_naiveOA = evidence_frame(reports, use_negation=True, use_oa_vocab=False).values.astype(float)\nS_full    = evidence_frame(reports, use_negation=True, use_oa_vocab=True ).values.astype(float)\nprint(\"osteoarthritis labels, naive literal reading vs consequence vocabulary:\")\nfor lab in [\"Medial OA\", \"Lateral OA\", \"PF OA\"]:\n    j = LABELS.index(lab); yj = y[:, j]\n    a0 = roc_auc_score(yj, S_naiveOA[:, j]); a1 = roc_auc_score(yj, S_full[:, j])\n    print(f\"  {lab:10s} naive AUC {a0:.3f}  ->  consequence AUC {a1:.3f}   (delta {a1-a0:+.3f})\")"},{"cell_type":"code","execution_count":null,"id":"6aea1827","metadata":{},"outputs":[],"source":"# FIG4  CENTERPIECE  OA resurrection\nfig, ax = plt.subplots(figsize=(8.5, 4.6))\noa = [\"Medial OA\", \"Lateral OA\", \"PF OA\"]\nnaive = [roc_auc_score(y[:, LABELS.index(l)], S_naiveOA[:, LABELS.index(l)]) for l in oa]\ncons  = [roc_auc_score(y[:, LABELS.index(l)], S_full[:, LABELS.index(l)])    for l in oa]\nx = np.arange(len(oa)); w = 0.36\nax.bar(x - w/2, naive, w, label=\"naive literal reading\", color=\"#b9c4cc\")\nax.bar(x + w/2, cons,  w, label=\"consequence vocabulary\", color=\"#c0603a\")\nax.axhline(0.5, ls=\"--\", color=\"k\", lw=1); ax.text(2.35, 0.51, \"chance\", fontsize=9)\nax.set_xticks(x); ax.set_xticklabels(oa); ax.set_ylabel(\"ROC AUC on 58 gold\"); ax.set_ylim(0, 1)\nfor xi, v in zip(x - w/2, naive): ax.text(xi, v + 0.01, f\"{v:.2f}\", ha=\"center\", fontsize=9)\nfor xi, v in zip(x + w/2, cons):  ax.text(xi, v + 0.01, f\"{v:.2f}\", ha=\"center\", fontsize=9)\nax.set_title(\"Osteoarthritis is almost never written as OA\")\nax.legend(loc=\"lower right\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"998aeae9","metadata":{},"source":"### 3b. Failure mode two: the word is present but negated\n\nReports state absence constantly: no tear, sin rotura, geen scheur, kein, pas de. A plain keyword hit\ncounts a ruled out finding as present. Sentence scope negation with cue lists in the report languages\nremoves those false positives. It is a real precision cleanup, and honestly a smaller mover on the macro\nAUC than the osteoarthritis fix."},{"cell_type":"code","execution_count":null,"id":"bea260c8","metadata":{},"outputs":[],"source":"def wilson(k, nn, z=1.96):\n    if nn == 0: return (float('nan'), float('nan'))\n    p = k/nn; d = 1 + z*z/nn\n    c = (p + z*z/(2*nn))/d; hw = z*math.sqrt(p*(1-p)/nn + z*z/(4*nn*nn))/d\n    return (max(0, c-hw), min(1, c+hw))\n\nS_noNeg = evidence_frame(reports, use_negation=False, use_oa_vocab=True).values.astype(float)\nprint(\"negation precision cleanup (threshold e>0, Wilson 95% CI on precision):\")\nprint(f\"{'label':18s} {'n+':>3s} | {'prec no-neg':>12s} {'prec neg':>10s} | {'rec no-neg':>10s} {'rec neg':>8s}\")\nneg_rows = []\nfor lab in [\"Fracture\",\"Baker's\",\"Medial Meniscus\",\"Effusion\",\"ACL\",\"Lateral Meniscus\"]:\n    j = LABELS.index(lab); yj = y[:, j]\n    out = {}\n    for tag, Smat in ((\"no\", S_noNeg), (\"neg\", S_full)):\n        pred = (Smat[:, j] > 0).astype(int)\n        tp = int(((pred==1)&(yj==1)).sum()); fp=int(((pred==1)&(yj==0)).sum()); fn=int(((pred==0)&(yj==1)).sum())\n        prec = tp/(tp+fp) if tp+fp else float('nan'); rec = tp/(tp+fn) if tp+fn else float('nan')\n        pl, ph = wilson(tp, tp+fp)\n        out[tag] = (prec, rec, pl, ph)\n    neg_rows.append((lab, out))\n    print(f\"{lab:18s} {int(yj.sum()):3d} |   {out['no'][0]:.2f}        {out['neg'][0]:.2f}    |   {out['no'][1]:.2f}      {out['neg'][1]:.2f}\")"},{"cell_type":"code","execution_count":null,"id":"c2e89342","metadata":{},"outputs":[],"source":"# FIG5  negation precision cleanup with Wilson CIs\nfig, ax = plt.subplots(figsize=(9, 4.6))\nlabs = [r[0] for r in neg_rows]\nx = np.arange(len(labs)); w = 0.36\np_no = [r[1]['no'][0] for r in neg_rows]; p_ne = [r[1]['neg'][0] for r in neg_rows]\ne_no = [[r[1]['no'][0]-r[1]['no'][2] for r in neg_rows], [r[1]['no'][3]-r[1]['no'][0] for r in neg_rows]]\ne_ne = [[r[1]['neg'][0]-r[1]['neg'][2] for r in neg_rows], [r[1]['neg'][3]-r[1]['neg'][0] for r in neg_rows]]\nax.bar(x - w/2, p_no, w, yerr=e_no, capsize=3, label=\"keyword presence\", color=\"#b9c4cc\")\nax.bar(x + w/2, p_ne, w, yerr=e_ne, capsize=3, label=\"with negation scope\", color=\"#3d7ea8\")\nax.set_xticks(x); ax.set_xticklabels(labs, rotation=20, ha=\"right\")\nax.set_ylabel(\"precision on 58 gold\"); ax.set_ylim(0, 1)\nax.set_title(\"Negation removes false positives (Wilson 95% CI)\")\nax.legend(loc=\"lower left\")\nplt.tight_layout(); plt.show()"},{"cell_type":"code","execution_count":null,"id":"35ee82e4","metadata":{},"outputs":[],"source":"# per language negation cue coverage: fraction of reports (per detected language) containing any neg cue\n# NEG_PRE and norm are already defined by the embedded labeler cell above\ncov = {}\ntmp = train.loc[has_report].copy()\ntmp[\"_lang\"] = lang.values\nfor lg in [l for l in tmp[\"_lang\"].unique() if l != \"Undetermined\"]:\n    sub = tmp[tmp[\"_lang\"] == lg][\"Report\"].apply(lambda t: any(nc in norm(t) for nc in NEG_PRE))\n    cov[lg] = sub.mean()\nprint(\"fraction of reports containing at least one negation cue, by detected language:\")\nfor k, v in sorted(cov.items(), key=lambda kv: -kv[1]):\n    print(f\"  {k:14s} {v:.2f}\")"},{"cell_type":"markdown","id":"c6185201","metadata":{},"source":"### 3c. The ablation ladder\n\nCredit assigned by the numbers, not by adjectives. Each rung is the macro AUC over all twelve labels on\nthe 58 gold. Negation is a genuine but modest mover; the consequence vocabulary step is the larger lever."},{"cell_type":"code","execution_count":null,"id":"e79fb373","metadata":{},"outputs":[],"source":"rungs = [(\"crude keyword presence\", False, False),\n         (\"plus sentence scope negation\", True, False),\n         (\"plus OA consequence vocabulary and tricompartmental\", True, True)]\nprev = None; macros = []\nfor name, neg, oa in rungs:\n    S = evidence_frame(reports, use_negation=neg, use_oa_vocab=oa).values.astype(float)\n    m, _ = macro_auc(y, S)\n    macros.append((name, m))\n    d = \"\" if prev is None else f\"   (delta {m-prev:+.3f})\"\n    print(f\"  {m:.3f}  {name}{d}\")\n    prev = m"},{"cell_type":"code","execution_count":null,"id":"db4bfd81","metadata":{},"outputs":[],"source":"# FIG6  ablation ladder\nfig, ax = plt.subplots(figsize=(8.5, 4.0))\nnames = [m[0] for m in macros]; vals = [m[1] for m in macros]\nax.plot(range(len(vals)), vals, \"-o\", color=\"#c0603a\", lw=2, ms=8)\nax.set_xticks(range(len(names)))\nax.set_xticklabels([\"crude\", \"+ negation\", \"+ OA vocab\"])\nfor i, v in enumerate(vals):\n    ax.text(i, v + 0.004, f\"{v:.3f}\", ha=\"center\", fontsize=10)\nax.set_ylabel(\"macro ROC AUC (58 gold)\")\nax.set_title(\"Ablation ladder: each rung adds one idea\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"0fca3829","metadata":{},"source":"## 4. Grade the labeler the way the metric grades a submission\n\nThe competition metric is the macro average of twelve ROC AUCs. Here is the labeler measured the same way,\nwith study level bootstrap confidence intervals so the tiny sample size is visible. Labels with fewer than\ntwelve positives are flagged sign only. This is an in sample estimate: the lexicon comes from external\nvocabulary but only 58 gold exist and there is no second labeled set, so treat these as optimistic. They\nare not a ceiling and not a leaderboard number."},{"cell_type":"code","execution_count":null,"id":"8bebebda","metadata":{},"outputs":[],"source":"m_full, per_full = macro_auc(y, S_full)\nrng = np.random.default_rng(20260809)\nB = 2000; n = len(gold)\nboot_macro = np.empty(B); boot_per = np.full((B, 12), np.nan)\nfor b in range(B):\n    idx = rng.integers(0, n, n)\n    mm, pp = macro_auc(y[idx], S_full[idx]); boot_macro[b] = mm; boot_per[b] = pp\ndef ci(a):\n    a = a[~np.isnan(a)]; return np.percentile(a, 2.5), np.percentile(a, 97.5)\nlo, hi = ci(boot_macro)\nprint(f\"macro AUC = {m_full:.3f}   95% study bootstrap CI [{lo:.3f}, {hi:.3f}]   B={B}\")\nprint(f\"\\n{'label':18s} {'n+':>3s} {'AUC':>6s} {'CI95':>16s} {'MCC':>6s} {'kappa':>6s}  flag\")\nrel_rows = []\nfor j, lab in enumerate(LABELS):\n    yj = y[:, j]; sj = S_full[:, j]; npj = int(yj.sum())\n    l, h = ci(boot_per[:, j]); pred = (sj > 0).astype(int)\n    mcc = matthews_corrcoef(yj, pred) if len(set(yj))>1 and len(set(pred))>1 else np.nan\n    kap = cohen_kappa_score(yj, pred) if len(set(yj))>1 else np.nan\n    flag = \"sign-only\" if npj < 12 else \"\"\n    rel_rows.append(dict(label=lab, n_pos=npj, auc=per_full[j], ci_lo=l, ci_hi=h, mcc=mcc, kappa=kap))\n    print(f\"{lab:18s} {npj:3d} {per_full[j]:6.3f}  [{l:5.3f},{h:5.3f}] {mcc:6.3f} {kap:6.3f}  {flag}\")"},{"cell_type":"code","execution_count":null,"id":"0351bd97","metadata":{},"outputs":[],"source":"# FIG7  per label AUC with bootstrap CI whiskers, sorted\nfig, ax = plt.subplots(figsize=(9, 5.2))\noi = np.argsort(per_full)\nyv = per_full[oi]\nlov = np.array([ci(boot_per[:, j])[0] for j in oi]); hiv = np.array([ci(boot_per[:, j])[1] for j in oi])\nerr = np.vstack([yv - lov, hiv - yv])\ncols = [\"#9aa7b1\" if int(y[:, j].sum()) < 12 else \"#3d7ea8\" for j in oi]\nax.errorbar(yv, range(12), xerr=err, fmt=\"o\", color=\"#c0603a\", ecolor=\"#888\", capsize=3, ms=7, ls=\"none\")\nfor i, j in enumerate(oi):\n    ax.scatter(yv[i], i, color=cols[i], zorder=3, s=45)\nax.axvline(0.5, ls=\"--\", color=\"k\", lw=1)\nax.set_yticks(range(12)); ax.set_yticklabels([f\"{LABELS[j]} (n+={int(y[:,j].sum())})\" for j in oi])\nax.set_xlabel(\"ROC AUC on 58 gold\"); ax.set_xlim(0.3, 1.0)\nax.set_title(f\"Per label labeler AUC with study level bootstrap 95% CI (macro {m_full:.3f})\")\nplt.tight_layout(); plt.show()\nprint(\"Blue points have n_pos >= 12; grey points are sign only (MCL, Lateral OA).\")"},{"cell_type":"code","execution_count":null,"id":"14220394","metadata":{},"outputs":[],"source":"# scrubbed error dump: a few false positives and false negatives, PHI safe (short snippet only)\ndef snippet(t, terms, width=70):\n    tn = t.lower()\n    for term in terms:\n        i = tn.find(term)\n        if i >= 0:\n            return \"...\" + t[max(0, i-width//2): i+width].replace(chr(10), \" \") + \"...\"\n    return t[:width].replace(chr(10), \" \") + \"...\"\nprint(\"sample disagreements (short scrubbed snippets):\")\nfor lab, terms in [(\"Fracture\", [\"fracture\",\"fractura\",\"fraktur\"]), (\"Effusion\", [\"effusion\",\"derrame\",\"erguss\"])]:\n    j = LABELS.index(lab); pred = (S_full[:, j] > 0).astype(int); yj = y[:, j]\n    fp = np.where((pred==1)&(yj==0))[0][:1]; fn = np.where((pred==0)&(yj==1))[0][:1]\n    for k in fp:\n        print(f\"  [{lab}] false positive: {snippet(reports[k], terms)}\")\n    for k in fn:\n        print(f\"  [{lab}] false negative: {snippet(reports[k], terms)}\")"},{"cell_type":"markdown","id":"f90647e7","metadata":{},"source":"## 5. Which plane and sequence each finding needs\n\nThis is the bridge to an images only model: route each finding to the series that actually shows it, using\nthe plane and fluid sensitive and fat suppression tags that exist in both train and test series metadata.\nThe map below is standard knee MRI reading knowledge, cross checked against what the metadata makes\navailable at test time."},{"cell_type":"code","execution_count":null,"id":"3a07e9ab","metadata":{},"outputs":[],"source":"# a compact finding -> plane/sequence table (reading knowledge), then availability check on test_series\nrouting = {\n \"ACL\":(\"Sagittal\",\"any\"), \"MCL\":(\"Coronal\",\"any\"),\n \"Medial Meniscus\":(\"Sagittal\",\"any\"), \"Lateral Meniscus\":(\"Sagittal\",\"any\"),\n \"Medial OA\":(\"Coronal\",\"any\"), \"Lateral OA\":(\"Coronal\",\"any\"), \"PF OA\":(\"Axial\",\"any\"),\n \"Effusion\":(\"Axial\",\"fluid sensitive\"), \"Synovitis\":(\"Axial\",\"fluid sensitive\"),\n \"Baker's\":(\"Axial\",\"fluid sensitive\"), \"Contusion\":(\"any\",\"fluid sensitive + fat suppressed\"),\n \"Fracture\":(\"any\",\"fluid sensitive + fat suppressed\"),\n}\ntest_planes = set(test_series[\"Anatomical_Plane\"].unique())\nprint(\"finding -> preferred plane / sequence  (available planes in test:\", sorted(test_planes), \")\")\nfor lab in LABELS:\n    pl, sq = routing[lab]\n    avail = \"yes\" if (pl == \"any\" or pl in test_planes) else \"MISSING\"\n    print(f\"  {lab:18s} {pl:10s} {sq:28s} plane available in test: {avail}\")\nfs = int((test_series[\"Fluid_Sensitive\"]==1).sum()); ffs = int(((test_series['Fluid_Sensitive']==1)&(test_series['Fat_Suppression']==1)).sum())\nprint(f\"\\ntest series that are fluid sensitive: {fs} of {len(test_series)}; fluid sensitive AND fat suppressed: {ffs}\")"},{"cell_type":"code","execution_count":null,"id":"b2541afa","metadata":{},"outputs":[],"source":"# FIG8  plane/sequence -> finding routing map\nfig, ax = plt.subplots(figsize=(9, 5))\nplanes3 = [\"Sagittal\",\"Coronal\",\"Axial\",\"any\"]\nseqs = [\"any\",\"fluid sensitive\",\"fluid sensitive + fat suppressed\"]\nM = np.zeros((12, len(planes3)))\nfor i, lab in enumerate(LABELS):\n    pl, sq = routing[lab]; M[i, planes3.index(pl)] = 1 + seqs.index(sq)*0.0 + (0.5 if sq!=\"any\" else 0)\nim = ax.imshow(M, cmap=\"Blues\", aspect=\"auto\")\nax.set_xticks(range(len(planes3))); ax.set_xticklabels(planes3)\nax.set_yticks(range(12)); ax.set_yticklabels(LABELS); ax.grid(False)\nfor i, lab in enumerate(LABELS):\n    pl, sq = routing[lab]\n    ax.text(planes3.index(pl), i, \"FS\" if \"fluid\" in sq else \"x\", ha=\"center\", va=\"center\",\n            fontsize=8, color=\"#08306b\")\nax.set_title(\"Preferred plane per finding (FS = needs a fluid sensitive sequence)\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"f4a823af","metadata":{},"source":"### 5a. The routing gallery: one real slice per finding\n\nThe matrix above, made concrete. For each of the 12 findings we deterministically select a gold study that\nis positive for it and show the central slice of the reading-convention plane. Labels are STUDY level: a tile\nis a study labeled positive for that finding, not a localized lesion on that slice. Every tile lists the\nother findings the same study carries (the 58 gold average over four positives each), so no tile pretends to\nbe a clean single-finding example. No arrows or boxes: nothing here is a verified pixel location."},{"cell_type":"code","execution_count":null,"id":"cc5edf4d","metadata":{},"outputs":[],"source":"# NEW-C  per-finding routing gallery: deterministic gold-example picker, provenance captions, no localization.\ntry:\n    done = False\n    if BASE_TRAIN and len(gold):\n        Gc = gold.sort_values(\"StudyInstanceUID\").reset_index(drop=True)\n        Yc = Gc[LABELS].fillna(0).astype(int).values\n        nposc = Yc.sum(0); totalc = Yc.sum(1)\n        ROUTE = {\"ACL\":(\"Sagittal\",False,False),\"MCL\":(\"Coronal\",False,False),\n                 \"Medial Meniscus\":(\"Sagittal\",False,False),\"Lateral Meniscus\":(\"Sagittal\",False,False),\n                 \"Medial OA\":(\"Coronal\",False,False),\"Lateral OA\":(\"Coronal\",False,False),\n                 \"PF OA\":(\"Axial\",False,False),\"Effusion\":(\"Axial\",True,False),\n                 \"Synovitis\":(\"Axial\",True,False),\"Baker's\":(\"Axial\",True,False),\n                 \"Contusion\":(\"any\",True,True),\"Fracture\":(\"any\",True,True)}\n        def passes(sid, plane, nf, nfs):\n            r = train_series[train_series.StudyInstanceUID == sid]\n            if plane != \"any\": r = r[r.Anatomical_Plane == plane]\n            if nf: r = r[r.Fluid_Sensitive == 1]\n            if nfs: r = r[r.Fat_Suppression == 1]\n            return len(r) > 0\n        fig, axes = plt.subplots(3, 4, figsize=(14, 11)); axes = axes.ravel()\n        tiers = {}\n        for li, lab in enumerate(LABELS):\n            ax = axes[li]; ax.set_axis_off()\n            plane, nf, nfs = ROUTE[lab]\n            pos = [i for i in range(len(Gc)) if Yc[i, li] == 1]\n            if not pos:\n                ax.text(0.5, 0.5, f\"no gold-positive\\nexample for {lab}\", ha=\"center\", va=\"center\", fontsize=9)\n                ax.set_title(f\"{lab}  n+={int(nposc[li])}  tier E\", fontsize=9); tiers[lab] = \"E\"; continue\n            pos.sort(key=lambda i: (0 if passes(Gc.StudyInstanceUID.iloc[i], plane, nf, nfs) else 1,\n                                    int(totalc[i] - 1), Gc.StudyInstanceUID.iloc[i]))\n            ci = pos[0]; sid = Gc.StudyInstanceUID.iloc[ci]\n            ok = passes(sid, plane, nf, nfs); burden = int(totalc[ci] - 1)\n            tier = \"A\" if (ok and burden == 0) else (\"B\" if (ok and burden <= 2) else (\"C\" if ok else \"D\"))\n            tiers[lab] = tier\n            rws = train_series[train_series.StudyInstanceUID == sid]\n            serid = pick_series_for(rws, plane, nf, nfs) or pick_series_for(rws, plane)                     or (rws.iloc[0].SeriesInstanceUID if len(rws) else None)\n            a = central_windowed(sdir(BASE_TRAIN, sid, serid)) if serid else None\n            copos = [LABELS[j] for j in range(12) if j != li and Yc[ci, j] == 1]\n            if a is not None:\n                ax.imshow(resize2d(a, 320), cmap=\"gray\", vmin=0, vmax=1)\n            else:\n                ax.text(0.5, 0.5, \"series unavailable\", ha=\"center\", va=\"center\", fontsize=9)\n            ax.set_title(f\"{lab}  n+={int(nposc[li])}  tier {tier}\", fontsize=9)\n            ax.text(4, 18, f\"{plane} FS{int(nf)} FatSat{int(nfs)}\", color=\"yellow\", fontsize=7)\n            ax.text(4, 312, \"also+: \" + (\", \".join(copos) if copos else \"none\"), color=\"deepskyblue\", fontsize=6)\n        fig.suptitle(\"Per-finding view gallery: each tile is a study gold-positive for that label at its reading plane. \"\n                     \"Study-level labels, not localized on the slice.\", fontsize=11)\n        plt.tight_layout(); plt.show()\n        print(\"achieved tiers:\", tiers)\n        print(\"Synovitis: on non-contrast MRI synovitis cannot be separated from a simple effusion; it is inferred, \"\n              \"not shown (would need post-gadolinium T1 fat-sat, absent here).\")\n        done = True\n    if not done: print(\"NEW-C: DICOM not available in this environment (renders on Kaggle)\")\nexcept Exception as e:\n    print(\"NEW-C skipped:\", repr(e))"},{"cell_type":"markdown","id":"234d2352","metadata":{},"source":"## 6. The payoff: silver labels for the 4349 report only studies\n\nRun the labeler on every report only study and emit soft per label scores plus an abstain flag for labels\nwhose structure is never mentioned. This is the artifact: training targets for any images only model, on\nstudies that shipped with no gold labels. Silver quality is bounded by the 58 gold validation above, not by\nground truth; it is a teacher signal, not an answer key."},{"cell_type":"code","execution_count":null,"id":"49215dd9","metadata":{},"outputs":[],"source":"report_only = train[~is_labeled & has_report].copy().reset_index(drop=True)\nS_ro = evidence_frame(report_only[\"Report\"].tolist(), use_negation=True, use_oa_vocab=True).values.astype(float)\n# soft probability via a global logistic squashing (ranking is unchanged; provides a 0..1 target)\na_glob, b_glob = 1.0, -0.5\nP_ro = 1.0 / (1.0 + np.exp(-(a_glob * S_ro + b_glob)))\nsilver = pd.DataFrame({\"StudyInstanceUID\": report_only[\"StudyInstanceUID\"].values})\nfor j, lab in enumerate(LABELS):\n    silver[lab] = P_ro[:, j].round(4)\nfor j, lab in enumerate(LABELS):\n    silver[lab + \"_abstain\"] = (S_ro[:, j] == 0).astype(int)\nsilver.to_csv(\"silver_labels.csv\", index=False)\n\nrel = pd.DataFrame(rel_rows)\nrel.to_csv(\"reliability_table.csv\", index=False)\nprint(\"wrote silver_labels.csv:\", silver.shape, \"and reliability_table.csv:\", rel.shape)\nprint(\"\\nper label silver positive count (soft score e>0) on the report only studies:\")\nsil_pos = (S_ro > 0).sum(axis=0)\nfor j, lab in enumerate(LABELS):\n    print(f\"  {lab:18s} gold+={int(y[:,j].sum()):3d}   silver+={int(sil_pos[j]):5d}   abstain={int((S_ro[:,j]==0).sum()):5d}\")"},{"cell_type":"code","execution_count":null,"id":"064c36aa","metadata":{},"outputs":[],"source":"# FIG9  silver yield: gold positives (58 studies) vs silver positives (4349 studies)\nfig, ax = plt.subplots(figsize=(9, 5.2))\noi = np.argsort(sil_pos)\ngp = np.array([int(y[:, j].sum()) for j in oi]); sp = sil_pos[oi]\nyy = np.arange(12)\nax.barh(yy, sp, color=\"#3d7ea8\", label=\"silver positives (of 4349)\")\nax.barh(yy, gp, color=\"#c0603a\", label=\"gold positives (of 58)\")\nax.set_yticks(yy); ax.set_yticklabels([LABELS[j] for j in oi])\nax.set_xlabel(\"positive studies\"); ax.legend(loc=\"lower right\")\nax.set_title(\"Training signal handed to any images only model\")\nfor i, v in enumerate(sp): ax.text(v + 10, i, str(int(v)), va=\"center\", fontsize=8)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"91542a48","metadata":{},"source":"## 7. How to fork this into an images only model\n\nThe report is train only, so the model that scores reads pixels. The recipe, deliberately not run here to\navoid a baseline that loses to the public leader on a budget this notebook cannot honestly validate:\n\n1. Load `silver_labels.csv` from this notebook output and join it to `train_series.csv` by study.\n2. Select one canonical series per plane, preferring fluid sensitive and fat suppressed for the bone and\n   fluid findings, using the plane and sequence map in section 5.\n3. Decode with pydicom once into a slice cache, invert MONOCHROME1, normalize per volume, and canonicalize\n   laterality from the DICOM Laterality tag so medial and lateral never flip. Do not use horizontal flip\n   augmentation, which swaps the medial and lateral compartments and scrambles four of the twelve labels.\n4. Train a per slice backbone with attention pooling across slices and planes and twelve sigmoid heads,\n   soft target BCE on the silver labels, with the 58 gold held out entirely as the only clean anchor.\n5. Attach any pretrained weights as a dataset and run inference images only, since the scored kernel has no\n   internet and no report text.\n\n```python\nsilver = pd.read_csv(\"/kaggle/input/<this-notebook-output>/silver_labels.csv\")\n# join to train_series, build the image pipeline above, train on silver, validate on the 58 gold\n```"},{"cell_type":"markdown","id":"e8cb4d70","metadata":{},"source":"## Limitations, stated plainly\n\n- The labeler AUCs are optimistic in sample estimates. The lexicon is external radiology vocabulary, but\n  only 58 gold exist and there is no second labeled set, so no number here is a ceiling or a leaderboard\n  preview.\n- The macro AUC is a train only label signal. It cannot be achieved at inference: the test set has no\n  reports, so any deployed model is images only.\n- MCL has nine positives and Lateral OA eleven, so both are flagged sign only; Effusion is intrinsically\n  hard and its AUC is low without being flagged; every per label number carries a bootstrap interval.\n- Silver labels are teacher labels whose quality is bounded by the 58 gold validation, not ground truth.\n- Language counts come from a crude offline detector with a large undetermined bucket, so the language\n  count is a lower bound.\n- Synovitis is weak without contrast; its low AUC is expected, not a defect to hide.\n- No leaderboard position is claimed. Other public work on this competition is not opened or benchmarked\n  here. This notebook is a train only labeling tool for any images only model, not a competing submission.\n\n## Sources\n\nCompanion notebooks on the same competition, each answering a different question:\n\n- [The RSNA knee labels are not independent](https://www.kaggle.com/code/busyaprime/the-rsna-knee-labels-are-not-independent) - a permutation test on the whole co-occurrence structure of the label set, against the default assumption of independent per label heads.\n- [The sequence name survived de-identification](https://www.kaggle.com/code/busyaprime/the-sequence-name-survived-de-identification) - which acquisition tags de-identification left in the DICOM headers, and how far they reconstruct Fluid_Sensitive, Fat_Suppression and Anatomical_Plane.\n- [Read the report then the knee silver baseline](https://www.kaggle.com/code/busyaprime/read-the-report-then-the-knee-silver-baseline) - the images only pipeline that consumes the silver labels written here and predicts the test set with no report available.\n"}],"metadata":{"kaggle":{"competitionSources":["rsna-knee-abnormality-detection"],"dataSources":[],"enableGpu":false,"enableInternet":false,"id":"busyaprime/osteoarthritis-is-almost-never-written-as-oa","isInternetEnabled":false,"language":"python","sourceType":"notebook","title":"Osteoarthritis is almost never written as OA"},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}