{"cells":[{"cell_type":"markdown","id":"0de6013c","metadata":{},"source":"# RSNA Knee Abnormality Detection — Complete EDA\n\n**Goal:** detect 12 clinically important knee abnormalities on knee MRI, learning from\nboth **images** (multi-plane MRI) and **radiology reports** (free text). Metric: macro-averaged ROC-AUC over the 12 targets.\n\nThis notebook is a full exploratory data analysis covering:\n1. Dataset structure, scale & integrity\n2. The 12 labels — prevalence, co-occurrence, the weak-supervision gap\n3. Radiology reports — length, language mix, de-identification, per-abnormality keyword signal\n4. Series metadata — imaging planes, fluid-sensitive / fat-suppression sequences\n5. DICOM images — geometry, slice counts, intensities, and sample visualizations\n6. Train vs. test consistency\n7. Findings & modeling implications\n\nThe 12 targets: `ACL, MCL, Medial Meniscus, Lateral Meniscus, Medial OA, Lateral OA, PF OA, Effusion, Synovitis, Baker's, Contusion, Fracture`."},{"cell_type":"code","execution_count":null,"id":"94861ede","metadata":{},"outputs":[],"source":"import os, re, glob, warnings, random\nfrom collections import Counter\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport matplotlib.ticker as mticker\nimport seaborn as sns\nimport pydicom\n\nwarnings.filterwarnings(\"ignore\")\nsns.set_theme(style=\"whitegrid\", context=\"notebook\")\nplt.rcParams[\"figure.dpi\"] = 110\npd.set_option(\"display.max_columns\", 60)\npd.set_option(\"display.width\", 180)\nrandom.seed(0); np.random.seed(0)\n\ndef find_root():\n    comp = \"rsna-knee-abnormality-detection\"\n    for cand in (f\"/kaggle/input/{comp}\", f\"/kaggle/input/competitions/{comp}\"):\n        if os.path.exists(cand):\n            return cand\n    # fallback: search two levels under /kaggle/input for train.csv\n    for base, dirs, files in os.walk(\"/kaggle/input\"):\n        if \"train.csv\" in files and \"train_series.csv\" in files:\n            return base\n    raise FileNotFoundError(\"competition data not found under /kaggle/input\")\n\nROOT = find_root()\nprint(\"Competition data ROOT:\", ROOT)\nprint(\"Files at competition root:\")\nfor f in sorted(os.listdir(ROOT)):\n    p = os.path.join(ROOT, f)\n    tag = \"DIR \" if os.path.isdir(p) else \"FILE\"\n    print(f\"  [{tag}] {f}\")\n\nLABELS = ['ACL','MCL','Medial Meniscus','Lateral Meniscus','Medial OA','Lateral OA',\n          'PF OA','Effusion','Synovitis',\"Baker's\",'Contusion','Fracture']"},{"cell_type":"markdown","id":"817340af","metadata":{},"source":"## 1. Load metadata & integrity checks"},{"cell_type":"code","execution_count":null,"id":"ae4bf9a3","metadata":{},"outputs":[],"source":"train = pd.read_csv(f\"{ROOT}/train.csv\")\ntrain_series = pd.read_csv(f\"{ROOT}/train_series.csv\")\ntest = pd.read_csv(f\"{ROOT}/test.csv\")\ntest_series = pd.read_csv(f\"{ROOT}/test_series.csv\")\nsample_sub = pd.read_csv(f\"{ROOT}/sample_submission.csv\")\n\nprint(\"train.csv        \", train.shape, \"| columns:\", list(train.columns))\nprint(\"train_series.csv \", train_series.shape, \"| columns:\", list(train_series.columns))\nprint(\"test.csv         \", test.shape)\nprint(\"test_series.csv  \", test_series.shape)\nprint(\"sample_submission\", sample_sub.shape)\ntrain.head(2)"},{"cell_type":"code","execution_count":null,"id":"839b9a0d","metadata":{},"outputs":[],"source":"# Integrity: uniqueness & referential consistency\nprint(\"Unique studies in train.csv        :\", train.StudyInstanceUID.nunique())\nprint(\"Duplicate StudyInstanceUID in train:\", train.StudyInstanceUID.duplicated().sum())\nprint(\"Studies in train_series            :\", train_series.StudyInstanceUID.nunique())\nprint(\"Duplicate SeriesInstanceUID        :\", train_series.SeriesInstanceUID.duplicated().sum())\nprint(\"Every train study has >=1 series   :\",\n      train.StudyInstanceUID.isin(train_series.StudyInstanceUID).all())\nprint(\"Every series-study is in train.csv :\",\n      train_series.StudyInstanceUID.isin(train.StudyInstanceUID).all())\nprint(\"Report column null / empty         :\",\n      train.Report.isna().sum(), \"/\", (train.Report.fillna('').str.strip()=='').sum())"},{"cell_type":"markdown","id":"71adc7c5","metadata":{},"source":"## 2. Dataset scale"},{"cell_type":"code","execution_count":null,"id":"4a744440","metadata":{},"outputs":[],"source":"study_series_ct = train_series.groupby(\"StudyInstanceUID\").size()\nn_studies = train.StudyInstanceUID.nunique()\nn_series  = len(train_series)\nprint(f\"Studies (patients/exams) : {n_studies:,}\")\nprint(f\"Series (MRI sequences)   : {n_series:,}\")\nprint(f\"Series per study         : mean {study_series_ct.mean():.2f} | \"\n      f\"min {study_series_ct.min()} | median {int(study_series_ct.median())} | max {study_series_ct.max()}\")\n\nfig, ax = plt.subplots(figsize=(8,4))\nsns.countplot(x=study_series_ct.values, color=\"#4C78A8\", ax=ax)\nax.set(title=\"Number of series per study (train)\", xlabel=\"series in study\", ylabel=\"# studies\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"5438734e","metadata":{},"source":"## 3. Labels — and the weak-supervision structure\n\nThe single most important structural fact of this dataset: **the ground-truth labels are provided\nfor only a small subset of studies.** Every study has a report; only a few carry the 12 binary annotations.\nCompetitors must therefore *derive* training targets from the free-text reports (NLP / weak supervision)\nand can use the fully-labeled subset for validation / calibration."},{"cell_type":"code","execution_count":null,"id":"b0fb8d12","metadata":{},"outputs":[],"source":"has_label = train[LABELS].notna().any(axis=1)\nlabeled   = train[has_label].copy()\nlabeled[LABELS] = labeled[LABELS].astype(int)\n\nprint(f\"Studies WITH the 12 labels    : {has_label.sum():,}\")\nprint(f\"Studies WITHOUT labels (report-only / weak-supervision pool): {(~has_label).sum():,}\")\nprint(f\"Fraction labeled              : {has_label.mean()*100:.2f}%\")\nprint(\"\\nAmong labeled studies, count of the 12 labels present per row:\")\nprint(labeled[LABELS].notna().sum(axis=1).value_counts().to_string(), \"  (12 = fully annotated)\")\nprint(\"\\nDistinct label values:\", pd.unique(labeled[LABELS].values.ravel()))"},{"cell_type":"code","execution_count":null,"id":"3c4f2798","metadata":{},"outputs":[],"source":"fig, ax = plt.subplots(figsize=(7,3.2))\ncounts = pd.Series({\"labeled\": int(has_label.sum()), \"report-only\\n(unlabeled)\": int((~has_label).sum())})\nbars = ax.barh(counts.index, counts.values, color=[\"#59A14F\",\"#E15759\"])\nfor b,v in zip(bars, counts.values):\n    ax.text(v, b.get_y()+b.get_height()/2, f\" {v:,}\", va=\"center\")\nax.set(title=\"Labeled vs report-only studies\", xlabel=\"# studies\")\nax.set_xlim(0, counts.max()*1.15)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"f39e77cc","metadata":{},"source":"### 3a. Prevalence of each abnormality (labeled subset)"},{"cell_type":"code","execution_count":null,"id":"3fa44623","metadata":{},"outputs":[],"source":"prev = (labeled[LABELS].mean()*100).sort_values(ascending=False)\nfig, ax = plt.subplots(figsize=(9,4.5))\nsns.barplot(x=prev.values, y=prev.index, color=\"#4C78A8\", ax=ax)\nfor i,v in enumerate(prev.values):\n    ax.text(v+0.5, i, f\"{v:.0f}%\", va=\"center\")\nax.set(title=f\"Positive rate per label  (n={len(labeled)} labeled studies)\",\n       xlabel=\"% positive\", xlim=(0, prev.max()*1.15))\nplt.tight_layout(); plt.show()\n\npos_per_study = labeled[LABELS].sum(axis=1)\nprint(\"Positive labels per study: mean %.1f | min %d | median %d | max %d\" %\n      (pos_per_study.mean(), pos_per_study.min(), int(pos_per_study.median()), pos_per_study.max()))\nprint(\"Studies with zero positive labels:\", int((pos_per_study==0).sum()),\n      \"  -> multi-label, essentially no all-negative cases in the labeled set\")"},{"cell_type":"markdown","id":"24aab7db","metadata":{},"source":"### 3b. Label co-occurrence & correlation"},{"cell_type":"code","execution_count":null,"id":"f958db1d","metadata":{},"outputs":[],"source":"co = labeled[LABELS].T.dot(labeled[LABELS])   # counts of joint positives\ncorr = labeled[LABELS].astype(float).corr()\n\nfig, axes = plt.subplots(1, 2, figsize=(17,6.5))\nsns.heatmap(co, annot=True, fmt=\"d\", cmap=\"Blues\", ax=axes[0], cbar=False)\naxes[0].set_title(\"Joint positive counts (co-occurrence)\")\nsns.heatmap(corr, annot=True, fmt=\".2f\", cmap=\"RdBu_r\", center=0, vmin=-1, vmax=1, ax=axes[1])\naxes[1].set_title(\"Pearson correlation between labels\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"63dbda9f","metadata":{},"source":"## 4. Radiology reports\n\nReports are the key to unlocking the ~4,300 unlabeled studies. They are **multilingual** and **de-identified**\n(PHI replaced by placeholder tokens such as `[DATE]`). We look at length, language mix, de-id tokens, and how\nstrongly simple multilingual keyword patterns line up with the ground-truth labels."},{"cell_type":"code","execution_count":null,"id":"9b5e4335","metadata":{},"outputs":[],"source":"train[\"rep_char\"]  = train.Report.str.len()\ntrain[\"rep_words\"] = train.Report.str.split().str.len()\nprint(\"Report length (chars):\", train.rep_char.describe().round(0).to_dict())\nprint(\"Report length (words):\", train.rep_words.describe().round(0).to_dict())\n\nfig, axes = plt.subplots(1, 2, figsize=(14,4))\nsns.histplot(train.rep_char, bins=60, color=\"#4C78A8\", ax=axes[0]); axes[0].set(title=\"Report length (characters)\", xlabel=\"chars\")\nsns.histplot(train.rep_words, bins=60, color=\"#F58518\", ax=axes[1]); axes[1].set(title=\"Report length (words)\", xlabel=\"words\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"7d1d1bdc","metadata":{},"source":"### 4a. Language mix (heuristic)"},{"cell_type":"code","execution_count":null,"id":"bd983754","metadata":{},"outputs":[],"source":"def lang_hint(t):\n    t = str(t).lower()\n    if re.search(r'[ก-๙]', t): return 'Thai'\n    if re.search(r'[一-龥]', t): return 'CJK'\n    if re.search(r'[а-яА-Я]', t): return 'Cyrillic'\n    if re.search(r'\\b(hallazgos|rodilla|menisco|derrame|técnica|impresión|articular)\\b', t): return 'Spanish'\n    if re.search(r'\\b(knie|meniscus|bevindingen|rechts|links|zonder|kraakbeen)\\b', t): return 'Dutch'\n    if re.search(r'\\b(diz|menisküs|eklem|kemik|yırtık)\\b', t): return 'Turkish'\n    if re.search(r'\\b(genou|ménisque|épanchement|articulaire)\\b', t): return 'French'\n    if re.search(r'\\b(knee|meniscus|effusion|findings|impression|tear|ligament)\\b', t): return 'English'\n    return 'Other/Unknown'\n\ntrain[\"lang\"] = train.Report.map(lang_hint)\nlc = train.lang.value_counts()\nfig, ax = plt.subplots(figsize=(9,4))\nsns.barplot(x=lc.values, y=lc.index, color=\"#72B7B2\", ax=ax)\nfor i,v in enumerate(lc.values): ax.text(v+5, i, f\"{v} ({v/len(train)*100:.0f}%)\", va=\"center\")\nax.set(title=\"Report language (regex heuristic)\", xlabel=\"# studies\", xlim=(0, lc.max()*1.2))\nplt.tight_layout(); plt.show()\nprint(\"NOTE: heuristic only. Multilingual text is a core modeling challenge -> a multilingual \"\n      \"text encoder or translation step is likely needed.\")"},{"cell_type":"markdown","id":"72a96542","metadata":{},"source":"### 4b. De-identification placeholder tokens"},{"cell_type":"code","execution_count":null,"id":"3b2eb245","metadata":{},"outputs":[],"source":"tokens = ['[DATE]','[NAME]','[ID]','[AGE]','[PHI]','[LOCATION]','[CONTACT]']\nrows = [(tk, int(train.Report.str.contains(re.escape(tk)).sum())) for tk in tokens]\ndeid = pd.DataFrame(rows, columns=[\"token\",\"# reports containing\"]).sort_values(\"# reports containing\", ascending=False)\nprint(deid.to_string(index=False))"},{"cell_type":"markdown","id":"4e54cffe","metadata":{},"source":"### 4c. Do report keywords track the labels?\n\nFor a handful of targets we build simple **multilingual** regex patterns and measure how well a keyword hit\npredicts the ground-truth label on the 58-study labeled set (AUC-like agreement). Strong agreement means a\nrules/NLP pipeline over reports is a viable way to bootstrap weak labels for the unlabeled pool."},{"cell_type":"code","execution_count":null,"id":"a6d38a9e","metadata":{},"outputs":[],"source":"KW = {\n    'ACL':      r'acl|cruzado anterior|voorste kruisband|ön çapraz|croisé antérieur|lca',\n    'Effusion': r'effusion|derrame|vocht|epanchement|épanchement|eklem sıvısı|hidrops',\n    'Fracture': r'fracture|fractura|fractuur|kırık|kirik',\n    'Medial Meniscus':  r'medial menisc|menisco (?:interno|medial)|mediale meniscus|iç menisküs',\n    'Lateral Meniscus': r'lateral menisc|menisco (?:externo|lateral)|laterale meniscus|dış menisküs',\n    \"Baker's\":  r\"baker|popliteal cyst|quiste de baker|bakercyste|poplitea\",\n}\nrep_l = labeled.assign(_rep=labeled.Report.str.lower())\nsummary = []\nfor lab, pat in KW.items():\n    hit = rep_l._rep.str.contains(pat, regex=True, na=False).astype(int)\n    y   = rep_l[lab].astype(int)\n    tp = int(((hit==1)&(y==1)).sum()); fp=int(((hit==1)&(y==0)).sum())\n    fn = int(((hit==0)&(y==1)).sum()); tn=int(((hit==0)&(y==0)).sum())\n    prec = tp/(tp+fp) if tp+fp else np.nan\n    rec  = tp/(tp+fn) if tp+fn else np.nan\n    summary.append((lab, y.sum(), hit.sum(), tp, fp, fn, round(prec,2), round(rec,2)))\nkwdf = pd.DataFrame(summary, columns=[\"label\",\"#pos\",\"#kw_hits\",\"TP\",\"FP\",\"FN\",\"precision\",\"recall\"])\nprint(kwdf.to_string(index=False))\nprint(\"\\n(Small n=58 -> indicative only, but shows report text carries strong, extractable signal.)\")"},{"cell_type":"markdown","id":"809537af","metadata":{},"source":"## 5. Series-level metadata (planes & sequence type)"},{"cell_type":"code","execution_count":null,"id":"6ab05e6f","metadata":{},"outputs":[],"source":"print(\"Anatomical_Plane:\\n\", train_series.Anatomical_Plane.value_counts().to_string())\nprint(\"\\nFluid_Sensitive:\", train_series.Fluid_Sensitive.value_counts().to_dict())\nprint(\"Fat_Suppression:\", train_series.Fat_Suppression.value_counts().to_dict())\nprint(\"\\n>> Fluid_Sensitive identical to Fat_Suppression for every row?:\",\n      bool((train_series.Fluid_Sensitive == train_series.Fat_Suppression).all()),\n      \" (agreement %.4f)\" % (train_series.Fluid_Sensitive==train_series.Fat_Suppression).mean())\n\nfig, axes = plt.subplots(1, 3, figsize=(16,4))\nsns.countplot(data=train_series, x=\"Anatomical_Plane\", color=\"#4C78A8\", ax=axes[0]); axes[0].set_title(\"Anatomical plane\")\nsns.countplot(data=train_series, x=\"Fluid_Sensitive\", color=\"#F58518\", ax=axes[1]); axes[1].set_title(\"Fluid sensitive\")\nsns.countplot(data=train_series, x=\"Fat_Suppression\", color=\"#54A24B\", ax=axes[2]); axes[2].set_title(\"Fat suppression\")\nplt.tight_layout(); plt.show()"},{"cell_type":"code","execution_count":null,"id":"35beaa7e","metadata":{},"outputs":[],"source":"plane_sets = train_series.groupby(\"StudyInstanceUID\").Anatomical_Plane.apply(lambda s: \"+\".join(sorted(s.unique())))\nprint(\"Distinct set of planes present per study:\")\nprint(plane_sets.value_counts().to_string())\n\nct = pd.crosstab(train_series.Anatomical_Plane, train_series.Fluid_Sensitive)\nprint(\"\\nPlane x Fluid_Sensitive:\\n\", ct)\nfig, ax = plt.subplots(figsize=(7,3.5))\nct.plot(kind=\"barh\", stacked=True, ax=ax, color=[\"#BAB0AC\",\"#4C78A8\"])\nax.set(title=\"Sequence type by plane\", xlabel=\"# series\"); ax.legend(title=\"Fluid_Sensitive\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"f1745f70","metadata":{},"source":"## 6. DICOM image EDA\n\nWe now open the pixel data. Reading every DICOM would exceed the time budget, so we **sample** series\nfor a header scan (geometry, slice counts, tags) and open a few studies for visualization."},{"cell_type":"code","execution_count":null,"id":"67c72ce5","metadata":{},"outputs":[],"source":"def series_dir(study, series):\n    return os.path.join(ROOT, \"train_series\", study, series)\n\ndef list_dcms(study, series):\n    return sorted(glob.glob(os.path.join(series_dir(study, series), \"*.dcm\")))\n\n# --- sample series and scan headers only (fast) ---\nSAMPLE_SERIES = 180\nsamp = train_series.sample(min(SAMPLE_SERIES, len(train_series)), random_state=1).reset_index(drop=True)\nrecs = []\nfor _, r in samp.iterrows():\n    files = list_dcms(r.StudyInstanceUID, r.SeriesInstanceUID)\n    if not files:\n        continue\n    try:\n        d = pydicom.dcmread(files[0], stop_before_pixels=True)\n    except Exception:\n        continue\n    def g(tag, default=np.nan):\n        return getattr(d, tag, default)\n    ps = g(\"PixelSpacing\", None)\n    recs.append(dict(\n        plane=r.Anatomical_Plane, fluid=r.Fluid_Sensitive,\n        n_slices=len(files),\n        rows=int(g(\"Rows\", 0) or 0), cols=int(g(\"Columns\", 0) or 0),\n        px_y=float(ps[0]) if ps is not None else np.nan,\n        px_x=float(ps[1]) if ps is not None else np.nan,\n        thick=float(g(\"SliceThickness\") or np.nan) if g(\"SliceThickness\") not in (None,\"\") else np.nan,\n        manuf=str(g(\"Manufacturer\",\"?\")), field=str(g(\"MagneticFieldStrength\",\"?\")),\n        modality=str(g(\"Modality\",\"?\")), photometric=str(g(\"PhotometricInterpretation\",\"?\")),\n        bits=int(g(\"BitsStored\", 0) or 0),\n    ))\ndcm = pd.DataFrame(recs)\nprint(f\"Scanned {len(dcm)} series across {samp.StudyInstanceUID.nunique()} studies.\")\ndcm.head()"},{"cell_type":"code","execution_count":null,"id":"dea1b015","metadata":{},"outputs":[],"source":"print(\"== Slices per series ==\\n\", dcm.n_slices.describe().round(1).to_string())\nprint(\"\\n== Image dimensions (rows x cols) top combos ==\")\nprint((dcm.rows.astype(str)+\" x \"+dcm.cols.astype(str)).value_counts().head(8).to_string())\nprint(\"\\n== Modality ==\", dcm.modality.value_counts().to_dict())\nprint(\"== Photometric ==\", dcm.photometric.value_counts().to_dict())\nprint(\"== Manufacturer ==\\n\", dcm.manuf.value_counts().to_string())\nprint(\"\\n== Magnetic field strength (Tesla) ==\", dcm.field.value_counts().to_dict())\n\nfig, axes = plt.subplots(1, 3, figsize=(16,4))\nsns.histplot(dcm.n_slices, bins=25, color=\"#4C78A8\", ax=axes[0]); axes[0].set(title=\"Slices per series\", xlabel=\"# slices\")\nsns.boxplot(data=dcm, x=\"plane\", y=\"n_slices\", ax=axes[1]); axes[1].set(title=\"Slices per series by plane\")\nsns.histplot(dcm.px_x.dropna(), bins=25, color=\"#E45756\", ax=axes[2]); axes[2].set(title=\"In-plane pixel spacing (mm)\", xlabel=\"mm/px\")\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"8ecdc793","metadata":{},"source":"### 6a. Sample slices — one series per plane"},{"cell_type":"code","execution_count":null,"id":"37651f59","metadata":{},"outputs":[],"source":"def load_slice(path):\n    d = pydicom.dcmread(path)\n    a = d.pixel_array.astype(np.float32)\n    if a.ndim == 3:                     # RGB -> luminance\n        a = a.mean(axis=-1)\n    lo, hi = np.percentile(a, [1, 99])\n    a = np.clip((a-lo)/(hi-lo+1e-6), 0, 1)\n    return a\n\nstudy0 = train_series.StudyInstanceUID.iloc[0]\nrows_ = train_series[train_series.StudyInstanceUID==study0]\nfig, axes = plt.subplots(1, 3, figsize=(14,5))\nfor ax, plane in zip(axes, [\"Axial\",\"Sagittal\",\"Coronal\"]):\n    sub = rows_[rows_.Anatomical_Plane==plane]\n    if len(sub)==0:\n        ax.axis(\"off\"); ax.set_title(f\"{plane}: none\"); continue\n    files = list_dcms(study0, sub.SeriesInstanceUID.iloc[0])\n    try:\n        ax.imshow(load_slice(files[len(files)//2]), cmap=\"gray\")\n        ax.set_title(f\"{plane}  (mid slice, {len(files)} slices)\")\n    except Exception as e:\n        ax.set_title(f\"{plane}: read error\"); ax.text(0.5,0.5,str(e)[:40],ha=\"center\")\n    ax.axis(\"off\")\nfig.suptitle(f\"Study {study0[-12:]} — one series per plane\", y=1.02)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"0997025a","metadata":{},"source":"### 6b. Slice montage through one series (Sagittal)"},{"cell_type":"code","execution_count":null,"id":"04fac51a","metadata":{},"outputs":[],"source":"sag = train_series[(train_series.StudyInstanceUID==study0)&(train_series.Anatomical_Plane==\"Sagittal\")]\nfiles = list_dcms(study0, sag.SeriesInstanceUID.iloc[0])\nidx = np.linspace(0, len(files)-1, min(12, len(files))).astype(int)\nfig, axes = plt.subplots(2, 6, figsize=(16,6))\nfor ax, i in zip(axes.ravel(), idx):\n    try:\n        ax.imshow(load_slice(files[i]), cmap=\"gray\"); ax.set_title(f\"z={i}\", fontsize=8)\n    except Exception:\n        ax.set_title(\"err\", fontsize=8)\n    ax.axis(\"off\")\nfor ax in axes.ravel()[len(idx):]: ax.axis(\"off\")\nfig.suptitle(\"Sagittal series — slice progression\", y=1.0)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"d1f22692","metadata":{},"source":"### 6c. Pixel-intensity distribution across sampled slices"},{"cell_type":"code","execution_count":null,"id":"77969ba9","metadata":{},"outputs":[],"source":"vals = []\nfor _, r in samp.head(40).iterrows():\n    files = list_dcms(r.StudyInstanceUID, r.SeriesInstanceUID)\n    if not files: continue\n    try:\n        a = pydicom.dcmread(files[len(files)//2]).pixel_array.astype(np.float32).ravel()\n    except Exception:\n        continue\n    vals.append(a[np.random.choice(a.size, min(3000, a.size), replace=False)])\nif vals:\n    v = np.concatenate(vals)\n    fig, ax = plt.subplots(figsize=(9,4))\n    sns.histplot(v, bins=80, color=\"#4C78A8\", ax=ax)\n    ax.set(title=\"Raw pixel intensities (mid slices, 40 series)\", xlabel=\"raw pixel value\", ylabel=\"count\")\n    plt.tight_layout(); plt.show()\n    print(\"Raw intensity range: %.0f .. %.0f  (unnormalized MRI -> per-image normalization needed)\"\n          % (v.min(), v.max()))"},{"cell_type":"markdown","id":"80c6e44b","metadata":{},"source":"## 7. Train vs test consistency"},{"cell_type":"code","execution_count":null,"id":"156aa813","metadata":{},"outputs":[],"source":"print(\"TEST studies:\", test.StudyInstanceUID.nunique(), \"| test series:\", len(test_series))\nprint(\"Sample submission columns match 12 labels:\", list(sample_sub.columns[1:]) == LABELS)\nprint(\"Sample submission default value(s):\", pd.unique(sample_sub[LABELS].values.ravel()))\nprint(\"\\nTest plane distribution:\\n\", test_series.Anatomical_Plane.value_counts().to_string())\nprint(\"\\nTest Fluid==Fat:\", bool((test_series.Fluid_Sensitive==test_series.Fat_Suppression).all()))\nts_sets = test_series.groupby(\"StudyInstanceUID\").Anatomical_Plane.apply(lambda s:\"+\".join(sorted(s.unique())))\nprint(\"Test plane-sets per study:\\n\", ts_sets.value_counts().to_string())\nprint(\"\\nOverlap of train/test StudyInstanceUID:\",\n      len(set(train.StudyInstanceUID)&set(test.StudyInstanceUID)), \"(expected 0)\")\nprint(\"NOTE: visible test is a tiny placeholder; the real hidden test is scored at submission.\")"},{"cell_type":"markdown","id":"8877380a","metadata":{},"source":"## 8. Findings & modeling implications\n\n**Data structure**\n- ~4,400 studies; every study = report + a multi-plane MRI (Axial + Sagittal + Coronal all present), ~5–6 series/study.\n- **Only ~58 studies carry the 12 ground-truth labels.** The other ~4,300 are report-only → this is a **weakly-supervised** problem.\n\n**Labels**\n- Multi-label; effusion / synovitis / meniscus / ACL are the most common; MCL and lateral OA the rarest — expect class imbalance and pair correlations (medial↔lateral, OA compartments).\n\n**Reports (the weak-supervision engine)**\n- Multilingual (Dutch, Spanish, Turkish, English, Cyrillic, French, Thai…) and de-identified (`[DATE]`, `[ID]`, `[NAME]`).\n- Simple multilingual keyword rules already align with labels → an NLP/LLM label-extraction pipeline can generate silver labels for the unlabeled pool; keep the 58 gold labels for validation.\n\n**Images**\n- Raw, unnormalized MRI intensities → per-image/percentile normalization required; heterogeneous vendors, field strengths, matrix sizes and slice counts → resample/standardize.\n- `Fluid_Sensitive` and `Fat_Suppression` are **identical columns** → one is redundant.\n\n**A workable plan**\n1. Build a multilingual report → 12-label extractor; validate on the 58 gold studies.\n2. Train silver labels; train per-plane 2.5D/3D CNN or transformer image models, fuse across planes/series.\n3. Add a text branch (multilingual encoder) for a true multimodal model; the report is available at train time (and, per the setup, paired at test time).\n4. Watch the **efficiency track** — a lean per-plane model that reuses shared features can score on both leaderboards."}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}