{"cells":[{"cell_type":"markdown","id":"4d3cdeb5","metadata":{},"source":"# RSNA Knee EDA: duplicate reports, a broken flag, and the real language mix\n\nTwelve binary findings per knee-MRI study — `ACL`, `MCL`, the two menisci, three\nosteoarthritis compartments, `Effusion`, `Synovitis`, `Baker's`, `Contusion`,\n`Fracture` — scored by macro-averaged AUC ROC.\n\nSeveral good EDAs already cover the ground you would expect: label prevalence, the\nco-occurrence matrix, plane and sequence counts, report length, sample DICOM slices.\n**This notebook deliberately does not repeat those plots.** It looks at the tabular\nfiles for the things that change what you build, and finds four:\n\n1. **`Fluid_Sensitive` and `Fat_Suppression` are byte-identical** across all 24,371\n   series. They describe different physics. One of them is not real data.\n2. **The quick substring trick for language ID is badly broken** on this corpus — the\n   version in circulation labels *zero* reports as Spanish while calling 422 Spanish\n   and 416 Turkish reports French. The real mix is seven Latin-script languages plus\n   Greek and Cyrillic.\n3. **Report text is not unique.** 46 report texts are shared by 177 studies, mostly\n   normal-study boilerplate — one identical template covers 37 studies. A random\n   split puts the same text in train and validation.\n4. **`Report` exists only in `train.csv`.** The test schema has no report column, so\n   the text is a training-time labelling signal, not a model input.\n\nTabular files only, no DICOM read, a few seconds on CPU."},{"cell_type":"code","execution_count":1,"id":"e44c7623","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:21.449332Z","iopub.status.busy":"2026-08-05T17:05:21.449218Z","iopub.status.idle":"2026-08-05T17:05:21.959146Z","shell.execute_reply":"2026-08-05T17:05:21.958184Z"}},"outputs":[],"source":"import re\nfrom collections import Counter\nfrom pathlib import Path\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\n\nKAGGLE_INPUT = Path(\"/kaggle/input\")\nCOMP = \"rsna-knee-abnormality-detection\"\n\n\ndef find_data_dir():\n    \"\"\"Locate the tabular files, on Kaggle or in a local checkout.\"\"\"\n    for p in [KAGGLE_INPUT / \"competitions\" / COMP, KAGGLE_INPUT / COMP,\n              Path(\"../data\"), Path(\"data\")]:\n        if (p / \"train.csv\").exists():\n            return p\n    # Kaggle mounts competition data in more than one layout. Search by the file\n    # we need, bounded to two levels so it never walks into the DICOM tree.\n    for pattern in (\"*/train.csv\", \"*/*/train.csv\"):\n        hits = sorted(KAGGLE_INPUT.glob(pattern))\n        if hits:\n            return hits[0].parent\n    return None\n\n\nDATA = find_data_dir()\nif DATA is None:\n    mounted = sorted(str(p) for p in KAGGLE_INPUT.glob(\"*\")) + \\\n              sorted(str(p) for p in KAGGLE_INPUT.glob(\"*/*\"))\n    raise FileNotFoundError(f\"train.csv not found. /kaggle/input holds: {mounted or 'nothing'}\")\nprint(f\"data directory: {DATA}\")\n\npd.set_option(\"display.width\", 200)\npd.set_option(\"display.max_columns\", 50)\n\ntrain = pd.read_csv(DATA / \"train.csv\")\nseries = pd.read_csv(DATA / \"train_series.csv\")\ntest = pd.read_csv(DATA / \"test.csv\")\ntest_series = pd.read_csv(DATA / \"test_series.csv\")\nsample_sub = pd.read_csv(DATA / \"sample_submission.csv\")\n\nLABELS = [c for c in sample_sub.columns if c != \"StudyInstanceUID\"]\nprint(f\"{len(LABELS)} targets: {', '.join(LABELS)}\")"},{"cell_type":"code","execution_count":2,"id":"81680274","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:21.960775Z","iopub.status.busy":"2026-08-05T17:05:21.960652Z","iopub.status.idle":"2026-08-05T17:05:21.964509Z","shell.execute_reply":"2026-08-05T17:05:21.963609Z"}},"outputs":[],"source":"# One visual language for the whole notebook. Surfaces are pinned light because\n# matplotlib output is a static PNG — it must stay legible under a dark UI theme.\nSURFACE, INK, INK_SOFT, GRID = \"#fcfcfb\", \"#0b0b0b\", \"#52514e\", \"#e6e5e1\"\nBLUE, ORANGE, MUTED = \"#2a78d6\", \"#eb6834\", \"#a9a8a2\"\n\nplt.rcParams.update({\n    \"figure.dpi\": 130,\n    \"figure.facecolor\": SURFACE,\n    \"savefig.facecolor\": SURFACE,\n    \"axes.facecolor\": SURFACE,\n    \"axes.edgecolor\": GRID,\n    \"axes.linewidth\": 0.8,\n    \"axes.labelcolor\": INK_SOFT,\n    \"axes.titlecolor\": INK,\n    \"axes.titlesize\": 11,\n    \"axes.titlelocation\": \"left\",\n    \"axes.titlepad\": 10,\n    \"axes.labelsize\": 9,\n    \"axes.spines.top\": False,\n    \"axes.spines.right\": False,\n    \"grid.color\": GRID,\n    \"grid.linewidth\": 0.8,\n    \"grid.linestyle\": \"-\",\n    \"text.color\": INK,\n    \"xtick.color\": INK_SOFT,\n    \"ytick.color\": INK_SOFT,\n    \"xtick.labelsize\": 9,\n    \"ytick.labelsize\": 9,\n    \"legend.frameon\": False,\n    \"legend.fontsize\": 9,\n    \"font.size\": 10,\n})\n\n\ndef grid_on(ax, axis=\"x\"):\n    \"\"\"Hairline grid on one axis only, behind the marks.\"\"\"\n    ax.set_axisbelow(True)\n    ax.grid(axis=axis)"},{"cell_type":"markdown","id":"932a396a","metadata":{},"source":"## 1. What the schemas actually contain\n\nTwo things here are worth more than a row count: a column the data description\npromises that does not exist, and a column that exists in train but not in test."},{"cell_type":"code","execution_count":3,"id":"20a70a33","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:21.965773Z","iopub.status.busy":"2026-08-05T17:05:21.965671Z","iopub.status.idle":"2026-08-05T17:05:21.992527Z","shell.execute_reply":"2026-08-05T17:05:21.991379Z"}},"outputs":[],"source":"shapes = pd.DataFrame(\n    [(\"train.csv\", *train.shape), (\"test.csv\", *test.shape),\n     (\"train_series.csv\", *series.shape), (\"test_series.csv\", *test_series.shape),\n     (\"sample_submission.csv\", *sample_sub.shape)],\n    columns=[\"file\", \"rows\", \"columns\"],\n).set_index(\"file\")\ndisplay(shapes.style.format(thousands=\",\"))"},{"cell_type":"code","execution_count":4,"id":"3c4d5084","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:21.993898Z","iopub.status.busy":"2026-08-05T17:05:21.993784Z","iopub.status.idle":"2026-08-05T17:05:22.002126Z","shell.execute_reply":"2026-08-05T17:05:22.001226Z"}},"outputs":[],"source":"# Which columns exist on each side of the train/test divide?\nschema = pd.DataFrame(index=sorted(set(train.columns) | set(test.columns)))\nschema[\"in_train\"] = [c in train.columns for c in schema.index]\nschema[\"in_test\"] = [c in test.columns for c in schema.index]\nschema[\"is_target\"] = [c in LABELS for c in schema.index]\ndisplay(schema.loc[~schema[\"is_target\"]].drop(columns=\"is_target\"))\n\nprint(f\"'Report'     in train: {'Report' in train.columns}   in test: {'Report' in test.columns}\")\nprint(f\"'PatientSex' in train: {'PatientSex' in train.columns}   \"\n      f\"(the data description lists it)\")"},{"cell_type":"markdown","id":"8a450a90","metadata":{},"source":"Two findings from that table.\n\n**`Report` is train-only.** Every training study carries free-text radiology, and the\ncompetition is described as multimodal, but the test schema has no report column. The\ntext is therefore a *labelling* signal for training, not an input your inference code\ncan consume. A fusion model with a text encoder has nothing to read at test time.\n\n**`PatientSex` is absent** despite appearing in the data description. There is also no\n`PatientID`, which leaves the grouping unit for a split unverified — section 3 comes\nback to that with actual evidence."},{"cell_type":"code","execution_count":5,"id":"ec7e7abd","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.003379Z","iopub.status.busy":"2026-08-05T17:05:22.003273Z","iopub.status.idle":"2026-08-05T17:05:22.009893Z","shell.execute_reply":"2026-08-05T17:05:22.009014Z"}},"outputs":[],"source":"labelled_mask = train[LABELS].notna().all(axis=1)\nn_labelled = int(labelled_mask.sum())\nlabelled = train[labelled_mask].copy()\n\nprint(f\"studies with all 12 labels : {n_labelled:>6,}  ({n_labelled / len(train):.1%})\")\nprint(f\"studies with report only   : {int((~labelled_mask).sum()):>6,}\")\nprint(f\"partially labelled studies : \"\n      f\"{int((train[LABELS].notna().any(axis=1) & ~labelled_mask).sum()):>6,}\")"},{"cell_type":"markdown","id":"daf9eddd","metadata":{},"source":"Labels are all-or-nothing — no study carries some of the twelve. Prevalence and\nco-occurrence over these 58 are covered well elsewhere, so here they are as plain\ntables rather than another pair of charts."},{"cell_type":"code","execution_count":6,"id":"359f7496","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.01109Z","iopub.status.busy":"2026-08-05T17:05:22.010988Z","iopub.status.idle":"2026-08-05T17:05:22.018156Z","shell.execute_reply":"2026-08-05T17:05:22.017248Z"}},"outputs":[],"source":"prev = (\n    labelled[LABELS]\n    .agg([\"sum\", \"mean\"])\n    .T.rename(columns={\"sum\": \"positives\", \"mean\": \"prevalence\"})\n    .sort_values(\"positives\", ascending=False)\n)\nprev[\"positives\"] = prev[\"positives\"].astype(int)\ndisplay(prev.style.format({\"prevalence\": \"{:.1%}\"}))"},{"cell_type":"code","execution_count":7,"id":"e11470d1","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.019348Z","iopub.status.busy":"2026-08-05T17:05:22.019242Z","iopub.status.idle":"2026-08-05T17:05:22.025669Z","shell.execute_reply":"2026-08-05T17:05:22.024837Z"}},"outputs":[],"source":"# Jaccard, P(both | either) — symmetric, so each pair appears once. Ranked instead\n# of drawn as a 12x12 matrix, which other EDAs already show.\nM = labelled[LABELS].values.astype(bool)\npairs = []\nfor i in range(len(LABELS)):\n    for j in range(i + 1, len(LABELS)):\n        both = int((M[:, i] & M[:, j]).sum())\n        either = int((M[:, i] | M[:, j]).sum())\n        pairs.append({\"finding A\": LABELS[i], \"finding B\": LABELS[j],\n                      \"both\": both, \"jaccard\": both / either if either else 0.0})\ntop_pairs = pd.DataFrame(pairs).sort_values(\"jaccard\", ascending=False).head(10)\ndisplay(top_pairs.style.format({\"jaccard\": \"{:.2f}\"}).hide(axis=\"index\"))"},{"cell_type":"markdown","id":"b52d4a17","metadata":{},"source":"## 2. The language mix, and why the usual shortcut gets it wrong\n\nThe reports are multilingual. The quick way to sort them is a cascade of substring\ntests — check for `'the '` and call it English, check for `'la '` and call it French,\nand so on. That approach is widely used and it fails hard here, for a reason worth\nseeing: **`'la '` is as common in Spanish as in French, and whichever language is\ntested first wins.**\n\nBelow, that cascade is run alongside a two-stage detector and the two are compared.\n\nStage one of the better version is Unicode script: if the letters are mostly Greek or\nCyrillic the *script* is settled with no ambiguity, though the language is not —\nCyrillic covers Bulgarian, Russian, Serbian and more, so those stay labelled as\nscripts. Stage two scores the Latin-script remainder by stopword counts across eight\ncandidate languages, and requires the winner to beat the runner-up by a clear margin.\nIt is still a heuristic; `unknown` is what it declines to call, reported rather than\nguessed away."},{"cell_type":"code","execution_count":8,"id":"f26c90d0","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.026928Z","iopub.status.busy":"2026-08-05T17:05:22.026826Z","iopub.status.idle":"2026-08-05T17:05:22.717441Z","shell.execute_reply":"2026-08-05T17:05:22.716515Z"}},"outputs":[],"source":"def detect_language_naive(text):\n    \"\"\"The substring cascade in common use. Order of the tests decides the answer.\"\"\"\n    t = text.lower()\n    if any(w in t for w in [\"the \", \"and \", \"of the\", \"is \", \"are \"]):\n        return \"English\"\n    if any(w in t for w in [\"la \", \"le \", \"et \", \"du \", \"des \"]):\n        return \"French\"\n    if any(w in t for w in [\"der \", \"die \", \"das \", \"und \", \"mit \"]):\n        return \"German\"\n    if any(w in t for w in [\" el \", \" la \", \" en \", \" de \", \" que \"]):\n        return \"Spanish\"\n    return \"unknown\"\n\n\n# Named as scripts, not languages, because that is all this stage can honestly claim.\nSCRIPTS = {\n    \"Greek script\": (0x0370, 0x03FF),\n    \"Cyrillic script\": (0x0400, 0x04FF),\n    \"Hebrew script\": (0x0590, 0x05FF),\n    \"Arabic script\": (0x0600, 0x06FF),\n    \"CJK\": (0x4E00, 0x9FFF),\n    \"Hangul\": (0xAC00, 0xD7AF),\n}\n\nSTOPWORDS = {\n    \"English\": \"the of and with is no in to are there without within left right knee normal\",\n    \"Spanish\": \"de la el en se con del no por una los las que sin rodilla menisco hallazgos\",\n    \"Portuguese\": \"de da do com em no na sem os as uma que dos joelho achados sinais\",\n    \"French\": \"de la le des en avec du les sans une aux genou articulaire aucune signal\",\n    \"Italian\": \"di la il del con in non una della dei ginocchio destro sinistro nella\",\n    \"German\": \"der die und mit im des ist ein kein von keine nachweis knie rechts links\",\n    \"Dutch\": \"van de het en met geen een is in bij niet knie rechts links bevindingen\",\n    \"Turkish\": \"ve bir ile mrg diz sag sol bulgular tetkik normaldir izlenmedi eklem kemik\",\n}\nSTOPWORDS = {k: set(v.split()) for k, v in STOPWORDS.items()}\n\nWORD_RE = re.compile(r\"[^\\W\\d_]+\", re.UNICODE)\n\n\ndef detect_language(text):\n    \"\"\"Unicode script first, then stopword scoring on Latin script.\"\"\"\n    letters = [c for c in text if c.isalpha()]\n    if not letters:\n        return \"unknown\"\n    for name, (lo, hi) in SCRIPTS.items():\n        if sum(lo <= ord(c) <= hi for c in letters) / len(letters) > 0.15:\n            return name\n\n    tokens = Counter(WORD_RE.findall(text.lower()))\n    scores = sorted(\n        ((sum(tokens[t] for t in stop), lang) for lang, stop in STOPWORDS.items()),\n        reverse=True,\n    )\n    (best, lang), (second, _) = scores[0], scores[1]\n    # Needs enough evidence, and a clear win over the runner-up: the Romance\n    # languages share function words, so a narrow margin is not a decision.\n    return lang if best >= 3 and best >= 1.3 * second else \"unknown\"\n\n\nreports = train[\"Report\"].dropna()\nlang = reports.apply(detect_language)\nlang_naive = reports.apply(detect_language_naive)\n\nmix = lang.value_counts()\nmix = mix.loc[[n for n in mix.index if n != \"unknown\"]\n              + ([\"unknown\"] if \"unknown\" in mix.index else [])]\ndisplay(pd.DataFrame({\"reports\": mix, \"share\": mix / len(reports)})\n        .style.format({\"reports\": \"{:,}\", \"share\": \"{:.1%}\"}))\n\nprint(\"what the substring cascade returns instead:\")\nprint(lang_naive.value_counts().to_string())"},{"cell_type":"markdown","id":"774de1c3","metadata":{},"source":"The cascade returns four buckets for a nine-bucket corpus, and one of its four —\nSpanish — never fires at all, because every Spanish report trips the French test\nfirst. The chart below breaks that down: for each language the two-stage detector\nidentifies, how often the cascade agrees, contradicts it, or gives up."},{"cell_type":"code","execution_count":9,"id":"20656d91","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.718943Z","iopub.status.busy":"2026-08-05T17:05:22.718833Z","iopub.status.idle":"2026-08-05T17:05:22.820688Z","shell.execute_reply":"2026-08-05T17:05:22.819698Z"}},"outputs":[],"source":"agree = pd.DataFrame({\"lang\": lang, \"naive\": lang_naive})\nrows = []\nfor name, g in agree.groupby(\"lang\"):\n    rows.append({\n        \"lang\": name,\n        \"agrees\": int((g[\"naive\"] == g[\"lang\"]).sum()),\n        \"contradicts\": int((g[\"naive\"] != g[\"lang\"]).sum() - (g[\"naive\"] == \"unknown\").sum()),\n        \"gives up\": int((g[\"naive\"] == \"unknown\").sum()),\n    })\ncmp_df = pd.DataFrame(rows).set_index(\"lang\")\ncmp_df = cmp_df.loc[cmp_df.sum(axis=1).sort_values().index]\n\nfig, ax = plt.subplots(figsize=(8, 4.2))\nleft = np.zeros(len(cmp_df))\nfor col, color in [(\"agrees\", BLUE), (\"contradicts\", ORANGE), (\"gives up\", MUTED)]:\n    ax.barh(cmp_df.index, cmp_df[col], left=left, height=0.6, color=color, label=col)\n    left += cmp_df[col].values\ngrid_on(ax, \"x\")\nax.set_xlabel(\"reports\")\nax.set_title(\"Substring cascade vs script + stopword detection\")\n# Below the axis: a three-entry legend collides with the title on the same line.\nax.legend(loc=\"upper center\", bbox_to_anchor=(0.5, -0.14), ncol=3)\nplt.tight_layout()\nplt.show()\n\nworst = cmp_df[\"contradicts\"].idxmax()\nprint(f\"most contradicted language: {worst} \"\n      f\"({cmp_df.loc[worst, 'contradicts']:,} reports)\")\nprint(f\"total reports the cascade gets wrong or abandons: \"\n      f\"{int(cmp_df['contradicts'].sum() + cmp_df['gives up'].sum()):,} of {len(reports):,}\")"},{"cell_type":"markdown","id":"b6eced82","metadata":{},"source":"The practical consequence: a report-based label extractor routed by the cascade sends\nSpanish and Turkish text down a French branch and never sees Greek or Cyrillic at all.\nThe failure is silent, and it correlates with imaging site rather than scattering at\nrandom — so it will not average out."},{"cell_type":"markdown","id":"753d6fff","metadata":{},"source":"## 3. Report text is not unique\n\nWorth checking before treating one report as one exam, or using text similarity to\nguess at the missing `PatientID`."},{"cell_type":"code","execution_count":10,"id":"9d911e2a","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.822144Z","iopub.status.busy":"2026-08-05T17:05:22.82203Z","iopub.status.idle":"2026-08-05T17:05:22.878577Z","shell.execute_reply":"2026-08-05T17:05:22.877474Z"}},"outputs":[],"source":"counts = reports.value_counts()\ndups = counts[counts > 1]\naffected = int(dups.sum())\n\nprint(f\"distinct report texts        : {len(counts):,}\")\nprint(f\"texts used by >1 study       : {len(dups):,}\")\nprint(f\"studies sharing text         : {affected:,} of {len(train):,} \"\n      f\"({affected / len(train):.1%})\")\nprint(f\"largest group                : {int(dups.max())} studies with identical text\")\nprint(f\"\\nmedian words, duplicated texts : {dups.index.to_series().str.split().str.len().median():.0f}\")\nprint(f\"median words, all reports      : {reports.str.split().str.len().median():.0f}\")"},{"cell_type":"code","execution_count":11,"id":"44184309","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.880184Z","iopub.status.busy":"2026-08-05T17:05:22.880067Z","iopub.status.idle":"2026-08-05T17:05:22.938248Z","shell.execute_reply":"2026-08-05T17:05:22.93736Z"}},"outputs":[],"source":"group_sizes = dups.value_counts().sort_index()\n\nfig, ax = plt.subplots(figsize=(8, 3.4))\nax.bar(group_sizes.index.astype(str), group_sizes.values, width=0.6, color=BLUE)\ngrid_on(ax, \"y\")\nax.set_xlabel(\"studies sharing one identical report text\")\nax.set_ylabel(\"report texts\")\nax.set_title(\"Duplicate report groups\")\nax.set_ylim(0, group_sizes.max() * 1.15)\nfor x, y in zip(range(len(group_sizes)), group_sizes.values):\n    ax.text(x, y + group_sizes.max() * 0.03, f\"{y}\", ha=\"center\", fontsize=9, color=INK_SOFT)\nplt.tight_layout()\nplt.show()"},{"cell_type":"code","execution_count":12,"id":"8f16c580","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.939631Z","iopub.status.busy":"2026-08-05T17:05:22.939519Z","iopub.status.idle":"2026-08-05T17:05:22.944739Z","shell.execute_reply":"2026-08-05T17:05:22.943861Z"}},"outputs":[],"source":"top_dups = pd.DataFrame({\n    \"studies\": dups.head(6).values,\n    \"words\": [len(t.split()) for t in dups.head(6).index],\n    \"language\": [detect_language(t) for t in dups.head(6).index],\n    \"text\": [t.strip().replace(\"\\n\", \" \")[:70] + \"...\" for t in dups.head(6).index],\n})\ndisplay(top_dups.style.hide(axis=\"index\"))"},{"cell_type":"markdown","id":"83ce8fbf","metadata":{},"source":"The duplicated texts are short — a median of 26 words against 129 across the corpus —\nand they read as normal-study templates: *\"Técnica: RMN de la rodilla. Resultados: Sin\nanomalías.\"* One such template covers 37 studies.\n\nTwo consequences.\n\n**Identical text does not mean the same patient.** These are boilerplate negatives, so\nreport text cannot substitute for the missing `PatientID`, and deduplicating on text\nwould throw away legitimately distinct studies.\n\n**A random split leaks anyway.** 177 studies share text with another study, so a\nrandom assignment puts byte-identical reports on both sides of the split. Any model\nreading the report memorises the pair rather than generalising, and the validation\nscore for those studies is free. Grouping on report text before splitting costs\nnothing and removes the effect."},{"cell_type":"markdown","id":"a518215b","metadata":{},"source":"## 4. `Fluid_Sensitive` and `Fat_Suppression` are the same column\n\n`train_series.csv` offers two sequence flags. They describe genuinely different\nphysics — T2/PD/STIR weighting versus fat suppression — and real protocols routinely\nhave one without the other, so they should disagree on a large share of series."},{"cell_type":"code","execution_count":13,"id":"4adf1692","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.945983Z","iopub.status.busy":"2026-08-05T17:05:22.945883Z","iopub.status.idle":"2026-08-05T17:05:22.94947Z","shell.execute_reply":"2026-08-05T17:05:22.948513Z"}},"outputs":[],"source":"differ = int((series[\"Fluid_Sensitive\"] != series[\"Fat_Suppression\"]).sum())\nprint(f\"series where Fluid_Sensitive != Fat_Suppression: {differ:,} of {len(series):,}\")\nprint(f\"identical on every row: {differ == 0}\")\n\nt_differ = int((test_series[\"Fluid_Sensitive\"] != test_series[\"Fat_Suppression\"]).sum())\nprint(f\"same check on test_series.csv: {t_differ:,} of {len(test_series):,} differ\")"},{"cell_type":"markdown","id":"3617d1fb","metadata":{},"source":"Zero disagreement on either side. This is almost certainly a defect in the released\nfiles rather than a property of knee MRI — two columns carrying one column's worth of\ninformation. Treat it as a single feature, expect a possible re-release, and do not\nbuild a plane-and-sequence encoding that assumes the two are independent.\n\nSince the flag is real information even if it is duplicated, here is the one view of\nit that is not just a marginal count — how sequence type splits across the three\nplanes:"},{"cell_type":"code","execution_count":14,"id":"1886b4d4","metadata":{"execution":{"iopub.execute_input":"2026-08-05T17:05:22.950717Z","iopub.status.busy":"2026-08-05T17:05:22.950616Z","iopub.status.idle":"2026-08-05T17:05:23.011283Z","shell.execute_reply":"2026-08-05T17:05:23.010353Z"}},"outputs":[],"source":"plane_by_flag = (\n    series.groupby([\"Anatomical_Plane\", \"Fluid_Sensitive\"]).size().unstack(fill_value=0)\n)\nshare = plane_by_flag.div(plane_by_flag.sum(axis=1), axis=0)\ndisplay(plane_by_flag.style.format(\"{:,}\"))\n\ncols = list(plane_by_flag.columns)\nx = np.arange(len(plane_by_flag.index))\nh = 0.36\nfig, ax = plt.subplots(figsize=(8, 3.6))\nfor k, (col, color) in enumerate(zip(cols, [BLUE, ORANGE])):\n    offset = (k - (len(cols) - 1) / 2) * (h + 0.04)\n    bars = ax.bar(x + offset, plane_by_flag[col], width=h, color=color,\n                  label=f\"Fluid_Sensitive = {col}\")\n    for bar, frac in zip(bars, share[col]):\n        ax.text(bar.get_x() + bar.get_width() / 2, bar.get_height() + 120,\n                f\"{frac:.0%}\", ha=\"center\", fontsize=8, color=INK_SOFT)\ngrid_on(ax, \"y\")\nax.set_xticks(x, plane_by_flag.index)\nax.set_ylabel(\"series\")\nax.set_title(\"Sequence type by plane  ·  labels are the within-plane share\")\nax.set_ylim(0, plane_by_flag.values.max() * 1.15)\nax.legend(loc=\"upper center\", bbox_to_anchor=(0.5, -0.16), ncol=2)\nplt.tight_layout()\nplt.show()"},{"cell_type":"markdown","id":"406d9cc8","metadata":{},"source":"Axial series are 80% fluid-sensitive while sagittal are barely half. So the flag is\nnot independent of plane, and a per-plane model sees a different sequence mix in each\nbranch — relevant if you plan separate heads per plane.\n\n## What this adds up to\n\n1. **Reports label, they do not predict.** `Report` is train-only, so the text is how\n   you get labels for the 4,349 unlabelled studies, and a text encoder has no input at\n   inference. Plan the pipeline as extraction-then-imaging, not fusion.\n2. **Extraction has to be genuinely multilingual.** Seven Latin-script languages plus\n   Greek and Cyrillic, and the substring cascade that looks like it handles this\n   silently routes two of the largest groups to the wrong branch.\n3. **Group on report text before splitting.** 177 studies share text with another\n   study. It is a one-line fix and it removes a real train/validation overlap.\n4. **Treat the two sequence flags as one.** They are byte-identical in train and test.\n5. **No `PatientID`, no `PatientSex`.** The grouping unit is unverified beyond the text\n   duplicates above, so whether one patient contributes several studies is still open.\n\n*Reproducible from the tabular files alone — no DICOM, no internet, a few seconds on\nCPU. Corrections and disagreement welcome in the comments.*"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.14.3"}},"nbformat":4,"nbformat_minor":5}