{"cells":[{"cell_type":"markdown","id":"f98d7928","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:10px solid #2a78d6;border-radius:8px;padding:26px 30px;margin:6px 0 20px 0;\"><div style=\"color:#c04b1c;font-size:12px;font-weight:700;letter-spacing:.16em;text-transform:uppercase;margin-bottom:10px;\">Exploratory data analysis</div><h1 style=\"color:#0b0b0b;font-size:31px;line-height:1.22;margin:0;font-weight:700;letter-spacing:-.01em;\">RSNA Knee Abnormality Detection: a first look</h1><div style=\"color:#6a6862;font-size:13px;margin-top:12px;\">Twelve findings per study &middot; macro-averaged ROC AUC &middot; 4,407 training studies, 58 labelled &middot; 265 GB</div></div>\n\nTwelve findings per knee MRI study, scored as the macro average of twelve ROC AUCs.\n4,407 training studies, 24,371 series, roughly 730,000 slices, 265 GB.\n\nTwo facts reshape the whole problem, and both are visible in the first ten minutes of\nlooking at the CSVs.\n\n**58 of the 4,407 training studies carry labels.** That is 1.3%. The other 4,349 studies\nship a free-text radiology report and nothing else. The competition is a label-extraction\nproblem wearing an imaging-competition costume, and the 58 gold studies are far too small\nand far too positive-enriched to rank two models against each other. Section 2 puts a\nnumber on that: a bootstrap 95% interval for macro AUC on the gold 58 spans about 0.079,\nwhich is an order of magnitude wider than the margins that separate leaderboard positions.\n\n**The reports exist at training time only.** `test.csv` has one column,\n`StudyInstanceUID`. There is no report at inference. Text is a label source and an\nauxiliary training signal; a pipeline that reads report text to predict scores nothing.\nSection 4 puts the two schemas one above the other.\n\nEverything below is computed from the competition data at run time. Nothing is\nhard-coded from a previous run. The notebook runs against the Kaggle mount and against a\npartial local mirror, and every DICOM walk is bounded by a constant declared in the config\ncell at the top of section 0.\n\n> **How this was made.** This notebook was written automatically by Claude Code running a\n> custom multi-agent LLM harness: parallel agents drafted the analysis, then separate\n> adversarial verifier agents executed every cell, checked every printed number, opened\n> every figure, and sent defects back to the builder script. It is published to Kaggle with\n> the Kaggle CLI, which is also how the companion baseline notebook is submitted. Full\n> detail is in the last section. **Code cells are collapsed by default** so this reads as a\n> report; click any *Show code* toggle to expand one."},{"cell_type":"markdown","id":"bcd9c903","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">0</span><span style=\"vertical-align:middle;\">Setup: input root, sampling budget, design system</span></h2></div>\n\nThree cells: find the data, declare every sampling bound in one place, then define the\nvisual system once so the rest of the notebook inherits it."},{"cell_type":"code","execution_count":null,"id":"728b5929","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Locate the competition data and declare every sampling bound.\"\"\"\n\nimport os\nimport re\nimport sys\nimport math\nimport json\nimport warnings\nimport collections\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\n\nwarnings.filterwarnings(\"ignore\")\n\n# ---- Sampling budget -------------------------------------------------------\n# Every DICOM walk in this notebook is bounded by one of these. Nothing here\n# iterates all ~730k slices; a Kaggle CPU session would not finish.\nSEED = 20260805                # every random draw in this notebook uses it\nN_HEADER_STUDIES = 90          # studies sampled for the DICOM header census\nMAX_HEADER_SERIES = 420        # hard cap on header reads (stop_before_pixels)\nMAX_FALLBACK_FILES = 400       # cap when the nested layout is absent (local runs)\nN_MONTAGE_SLICES = 4           # slices shown per plane in the pixel montage\nN_MONTAGE_CANDIDATES = 8       # series tried per plane before giving up\nN_INTENSITY_SERIES = 6         # series in the intensity-distribution panel\nN_BOOTSTRAP = 2000             # bootstrap resamples for the AUC noise floor\n\nLABELS = [\n    \"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n    \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\", \"Contusion\",\n    \"Fracture\",\n]\n\nRNG = np.random.default_rng(SEED)\n\n\ndef find_root():\n    \"\"\"Locate the directory holding ``train.csv``, Kaggle mount first.\n\n    Returns:\n        A ``Path`` to the data root.\n\n    Raises:\n        FileNotFoundError: If no candidate directory contains ``train.csv``.\n    \"\"\"\n    here = Path.cwd()\n    candidates = [\n        Path(\"/kaggle/input/rsna-knee-abnormality-detection\"),\n        Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\"),\n    ]\n    for up in [here, *here.parents][:5]:\n        candidates.append(up / \"data\" / \"kaggle_mirror\" / \"rsna-knee-abnormality-detection\")\n        candidates.append(up / \"data\" / \"raw\")\n    for c in candidates:\n        if (c / \"train.csv\").is_file():\n            return c\n    # Last resort: a shallow scan of /kaggle/input.\n    base = Path(\"/kaggle/input\")\n    if base.is_dir():\n        for child in sorted(base.iterdir()):\n            if (child / \"train.csv\").is_file():\n                return child\n    raise FileNotFoundError(\n        \"Could not locate train.csv. Checked: \" + \", \".join(str(c) for c in candidates)\n    )\n\n\nROOT = find_root()\nON_KAGGLE = str(ROOT).startswith(\"/kaggle/\")\n\n# Where DICOMs live, if they are here at all.\nTRAIN_DCM_DIR = ROOT / \"train_series\"\nTEST_DCM_DIR = ROOT / \"test_series\"\n\n# A local checkout keeps a handful of real DICOMs outside the mirror so this\n# notebook can be exercised offline. Harmless and absent on Kaggle.\nEXTRA_DCM_DIRS = [p for p in\n                  [ROOT.parent / \"sample\", ROOT.parent.parent / \"data\" / \"sample\"]\n                  if p.is_dir()]\n\nprint(\"data root      :\", ROOT)\nprint(\"running on     :\", \"Kaggle\" if ON_KAGGLE else \"local mirror\")\nprint(\"train DICOM dir:\", TRAIN_DCM_DIR, \"->\", \"present\" if TRAIN_DCM_DIR.is_dir() else \"ABSENT\")\nprint(\"test DICOM dir :\", TEST_DCM_DIR, \"->\", \"present\" if TEST_DCM_DIR.is_dir() else \"ABSENT\")\nif EXTRA_DCM_DIRS:\n    print(\"extra DICOM dir:\", \", \".join(str(p) for p in EXTRA_DCM_DIRS))\nprint(\"csv files      :\", sorted(p.name for p in ROOT.glob(\"*.csv\")))"},{"cell_type":"code","execution_count":null,"id":"e945f6f3","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Define the design system once. Every later chart inherits it.\"\"\"\n\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nfrom matplotlib.colors import LinearSegmentedColormap\nfrom matplotlib.lines import Line2D\nimport matplotlib.patheffects as pe\n\n# ---- Ink and surface -------------------------------------------------------\n# Three text tiers, every one of them at or above the WCAG AA 4.5:1 floor\n# against SURFACE. Ratios are computed, not eyeballed: 19.2, 7.7, 5.4. The two\n# structural greys below carry no text and are deliberately left light, because\n# darkening a gridline adds noise and contrast rules do not apply to it.\nSURFACE = \"#fcfcfb\"   # chart surface\nINK = \"#0b0b0b\"       # primary text                    19.2:1\nINK_2 = \"#52514e\"     # secondary text, value annotations 7.7:1\nMUTED = \"#6a6862\"     # axis labels, tick labels, captions 5.4:1\nGRID = \"#e1e0d9\"      # hairline gridlines, non-text\nAXIS = \"#c3c2b7\"      # baseline, non-text\n\n# ---- Categorical palette ---------------------------------------------------\n# Five hues, one per anatomical family. Chosen by exhaustive search over the\n# eight-slot source palette and validated all-pairs: worst pair CVD dE 13.0,\n# worst pair normal-vision dE 16.3 (OKLab x100). Twelve separate hues would\n# fail every colourblind-safety check, so colour carries the FAMILY and the\n# finding name is always written on the axis. No information is colour-alone.\nFAMILY_COLORS = {\n    \"Ligament\": \"#2a78d6\",            # blue\n    \"Meniscus\": \"#eda100\",            # yellow\n    \"Osteoarthritis\": \"#e87ba4\",      # magenta\n    \"Effusion & synovium\": \"#008300\",  # green\n    \"Bone trauma\": \"#4a3aa7\",         # violet\n}\nLABEL_FAMILY = {\n    \"ACL\": \"Ligament\", \"MCL\": \"Ligament\",\n    \"Medial Meniscus\": \"Meniscus\", \"Lateral Meniscus\": \"Meniscus\",\n    \"Medial OA\": \"Osteoarthritis\", \"Lateral OA\": \"Osteoarthritis\",\n    \"PF OA\": \"Osteoarthritis\",\n    \"Effusion\": \"Effusion & synovium\", \"Synovitis\": \"Effusion & synovium\",\n    \"Baker's\": \"Effusion & synovium\",\n    \"Contusion\": \"Bone trauma\", \"Fracture\": \"Bone trauma\",\n}\nLABEL_COLORS = {k: FAMILY_COLORS[v] for k, v in LABEL_FAMILY.items()}\n\nPRIMARY = \"#2a78d6\"    # the single-series colour, marks only\nACCENT = \"#eb6834\"     # emphasis: the one thing the chart is about, marks only\nNEUTRAL = \"#c3c2b7\"    # de-emphasised bars\n\n# Text-safe partners for the two mark colours. A mark does not have to clear a\n# contrast floor; the sentence next to it does. ACCENT at 3.1:1 is fine as a\n# bar and unreadable as a caption, so annotations that belong to an accent mark\n# are drawn in ACCENT_TEXT (4.8:1) and fills that carry SURFACE-coloured text\n# use PRIMARY_DEEP (5.3:1) rather than PRIMARY (4.3:1).\nACCENT_TEXT = \"#c04b1c\"   # accent annotations           4.8:1 on SURFACE\nPRIMARY_DEEP = \"#256abf\"  # fill behind SURFACE text     5.3:1 either way\n\n# ---- Sequential ramp (one hue, light to dark) ------------------------------\nBLUE_STEPS = [\"#cde2fb\", \"#b7d3f6\", \"#9ec5f4\", \"#86b6ef\", \"#6da7ec\", \"#5598e7\",\n              \"#3987e5\", \"#2a78d6\", \"#256abf\", \"#1c5cab\", \"#184f95\", \"#104281\",\n              \"#0d366b\"]\nSEQ = LinearSegmentedColormap.from_list(\"knee_blue\", BLUE_STEPS)\n\n\n# ---- Contrast ---------------------------------------------------------------\ndef _rel_lum(colour):\n    \"\"\"WCAG relative luminance of a colour.\n\n    Args:\n        colour: Anything matplotlib can parse as a colour.\n\n    Returns:\n        Relative luminance in ``[0, 1]``.\n    \"\"\"\n    total = 0.0\n    for weight, chan in zip((0.2126, 0.7152, 0.0722), mpl.colors.to_rgb(colour)):\n        lin = chan / 12.92 if chan <= 0.03928 else ((chan + 0.055) / 1.055) ** 2.4\n        total += weight * lin\n    return total\n\n\ndef contrast(fg, bg):\n    \"\"\"WCAG 2.x contrast ratio between two colours.\n\n    Args:\n        fg: Foreground colour.\n        bg: Background colour.\n\n    Returns:\n        A ratio between 1.0 and 21.0. AA wants 4.5 for normal-size text.\n    \"\"\"\n    lo, hi = sorted((_rel_lum(fg), _rel_lum(bg)))\n    return (hi + 0.05) / (lo + 0.05)\n\n\ndef ink_on(fill):\n    \"\"\"Pick whichever of INK and SURFACE reads better on a filled mark.\n\n    A number written inside a bar or a heatmap cell sits on the mark, not on\n    the page, so the page's three text tiers do not apply to it. Choosing by\n    computed contrast beats a hand-tuned lightness threshold, which is always\n    wrong somewhere in the middle of a sequential ramp.\n\n    Args:\n        fill: The colour of the mark the text sits on.\n\n    Returns:\n        ``INK`` or ``SURFACE``.\n    \"\"\"\n    return INK if contrast(INK, fill) >= contrast(SURFACE, fill) else SURFACE\n\n\ndef _seq_safe(t):\n    \"\"\"Sample the sequential ramp, stepping over its unreadable stretch.\n\n    About 58% of the way up the ramp the blue is mid-toned enough that neither\n    INK nor SURFACE clears 4.5:1 on it, so a number written in a cell of that\n    colour cannot be made readable by choosing the text colour. The band is 4%\n    of the ramp wide; nudging a value to the nearer edge is invisible in the\n    shading and makes every annotated cell legible.\n\n    Args:\n        t: Position on the ramp in ``[0, 1]``.\n\n    Returns:\n        An RGBA tuple.\n    \"\"\"\n    lo, hi = 0.560, 0.620\n    if lo < t < hi:\n        t = lo if t < (lo + hi) / 2 else hi\n    return SEQ(t)\n\n\n# The ramp actually used for filled cells that carry a number.\nSEQ_SAFE = mpl.colors.ListedColormap(\n    [_seq_safe(i / 255.0) for i in range(256)], name=\"knee_blue_safe\")\n\n\nmpl.rcParams.update({\n    \"figure.facecolor\": SURFACE,\n    \"figure.dpi\": 118,\n    \"savefig.facecolor\": SURFACE,\n    \"axes.facecolor\": SURFACE,\n    \"axes.edgecolor\": AXIS,\n    \"axes.linewidth\": 0.9,\n    \"axes.spines.top\": False,\n    \"axes.spines.right\": False,\n    \"axes.labelcolor\": MUTED,\n    \"axes.labelsize\": 12.5,\n    \"axes.titlesize\": 13.5,\n    \"axes.grid\": False,\n    \"grid.color\": GRID,\n    \"grid.linewidth\": 0.8,\n    \"grid.linestyle\": \"-\",\n    \"xtick.color\": MUTED,\n    \"ytick.color\": MUTED,\n    \"xtick.labelsize\": 12,\n    \"ytick.labelsize\": 12,\n    \"xtick.major.size\": 0,\n    \"ytick.major.size\": 0,\n    \"legend.frameon\": False,\n    \"legend.fontsize\": 12,\n    \"lines.linewidth\": 2.4,\n    \"lines.markersize\": 9,\n    \"font.family\": \"sans-serif\",\n    \"font.sans-serif\": [\"DejaVu Sans\", \"Segoe UI\", \"Helvetica\", \"Arial\"],\n    \"font.size\": 12.5,\n    \"text.color\": INK,\n})\n\n\ndef style(ax, grid_axis=\"x\"):\n    \"\"\"Apply the recessive grid and spine treatment to one axes.\n\n    Args:\n        ax: A matplotlib axes.\n        grid_axis: ``\"x\"``, ``\"y\"`` or ``\"none\"``. Gridlines run along this axis\n            only, behind the marks.\n\n    Returns:\n        The same axes, for chaining.\n    \"\"\"\n    for side in (\"top\", \"right\"):\n        ax.spines[side].set_visible(False)\n    for side in (\"left\", \"bottom\"):\n        ax.spines[side].set_color(AXIS)\n    ax.grid(False)\n    if grid_axis == \"x\":\n        ax.xaxis.grid(True, zorder=0)\n    elif grid_axis == \"y\":\n        ax.yaxis.grid(True, zorder=0)\n    ax.set_axisbelow(True)\n    ax.tick_params(length=0)\n    return ax\n\n\ndef _wrap(text, width_in, fontsize):\n    \"\"\"Wrap a header or caption string to the figure's own width.\n\n    The notebook's inline backend saves with a tight bounding box, so a text\n    run wider than the figure silently widens the exported PNG and leaves the\n    plot stranded in one corner. Wrapping to the figure width prevents that.\n\n    Args:\n        text: The string to wrap.\n        width_in: Figure width in inches.\n        fontsize: Point size the string will be drawn at.\n\n    Returns:\n        The string with newlines inserted.\n    \"\"\"\n    import textwrap\n    chars = max(24, int((width_in - 0.15) * 72.0 / (fontsize * 0.53)))\n    return \"\\n\".join(textwrap.wrap(text, chars)) if text else \"\"\n\n\ndef finish(fig, finding, detail=\"\", caption=\"\", tight=True, headroom=0.0,\n           footroom=0.0):\n    \"\"\"Add the finding title, optional detail line and caption, then show.\n\n    Titles state the finding, not the mechanism. All reserved space is computed\n    in inches from the wrapped line count, so a tall figure and a short one get\n    the same visual header band.\n\n    Args:\n        fig: The figure to close out.\n        finding: The headline. One sentence, stating what the chart shows.\n        detail: Optional second line in secondary ink.\n        caption: Optional footer in muted ink, for provenance and caveats.\n        tight: Run ``tight_layout``. Set ``False`` for image grids that manage\n            their own spacing.\n        headroom: Extra inches reserved above the axes, for figures whose\n            panels carry their own titles.\n        footroom: Extra inches reserved below the axes, on top of the measured\n            clearance, for anything the tight bounding box does not see.\n    \"\"\"\n    w, h = fig.get_size_inches()\n    finding = _wrap(finding, w, 17.5)\n    detail = _wrap(detail, w, 13.0)\n    caption = _wrap(caption, w, 11.25)\n    n_find = finding.count(\"\\n\") + 1\n    n_det = (detail.count(\"\\n\") + 1) if detail else 0\n    n_cap = (caption.count(\"\\n\") + 1) if caption else 0\n\n    head_in = 0.32 + 0.31 * n_find + (0.10 + 0.25 * n_det) + headroom\n    foot_in = ((0.18 + 0.21 * n_cap) if caption else 0.13) + footroom\n\n    # tight_layout silently does nothing on two whole classes of figure and only\n    # emits a UserWarning: any figure holding a fixed-aspect axes (imshow,\n    # set_aspect(\"equal\")), and any figure whose gridspec was given an explicit\n    # hspace or wspace, which is every stacked pair in this notebook. So the\n    # header band cannot be left to it. Run it for the tick-label margins, then\n    # clamp top and bottom directly, which always applies.\n    if tight:\n        try:\n            fig.tight_layout(rect=(0.0, foot_in / h, 1.0, 1.0 - head_in / h))\n        except Exception:\n            pass\n    sp = fig.subplotpars\n    top = min(sp.top, 1.0 - head_in / h)\n    bottom = max(sp.bottom, foot_in / h)\n    if top > bottom + 0.05:\n        fig.subplots_adjust(top=top, bottom=bottom)\n\n    # When tight_layout declined, the axes box moved but its tick labels and\n    # axis label still hang below it, unmeasured, and the caption is pinned to\n    # the figure edge. Measure the real overhang and lift the axes if the two\n    # would meet. Only ever lifts, never lowers, and never past a sane cap.\n    if caption:\n        try:\n            fig.canvas.draw()\n            rend = fig.canvas.get_renderer()\n            low = min(a.get_tightbbox(rend).y0 for a in fig.axes) / fig.bbox.height\n            shift = min(foot_in / h - low, 0.25)\n            sp = fig.subplotpars\n            if shift > 0.002 and sp.top > sp.bottom + shift + 0.05:\n                fig.subplots_adjust(bottom=sp.bottom + shift)\n        except Exception:\n            pass\n\n    fig.text(0.008, 1.0 - 0.20 / h, finding, ha=\"left\", va=\"top\",\n             fontsize=17.5, fontweight=\"bold\", color=INK, linespacing=1.25)\n    if detail:\n        fig.text(0.008, 1.0 - (0.25 + 0.31 * n_find) / h, detail, ha=\"left\",\n                 va=\"top\", fontsize=13, color=INK_2, linespacing=1.3)\n    if caption:\n        fig.text(0.008, 0.09 / h, caption, ha=\"left\", va=\"bottom\",\n                 fontsize=11.25, color=MUTED, linespacing=1.3)\n    plt.show()\n\n\ndef hbar(ax, names, values, colors=None, fmt=\"{:,.0f}\", annot=None,\n         pad_frac=0.022, xpad=1.18):\n    \"\"\"Draw a horizontal bar chart with every bar directly labelled.\n\n    Horizontal bars sorted by value beat vertical bars with rotated ticks, and a\n    value written at the bar end beats making the reader trace back to an axis.\n    Annotations wear secondary ink, never the series colour.\n\n    Args:\n        ax: Target axes.\n        names: Category names, drawn top to bottom in the order given.\n        values: Numeric values, same length as ``names``.\n        colors: Per-bar colours, or ``None`` for the single-series colour.\n        fmt: Format string applied to each value when ``annot`` is ``None``.\n        annot: Pre-formatted annotation strings, one per bar, overriding ``fmt``.\n        pad_frac: Annotation offset as a fraction of the axis range.\n        xpad: Right-hand headroom multiplier, so annotations never clip.\n\n    Returns:\n        The bar container.\n    \"\"\"\n    names = list(names)\n    values = np.asarray(values, dtype=float)\n    y = np.arange(len(names))[::-1]\n    if colors is None:\n        colors = [PRIMARY] * len(names)\n    if annot is None:\n        annot = [fmt.format(v) for v in values]\n    bars = ax.barh(y, values, height=0.70, color=colors, zorder=3)\n    ax.set_yticks(y)\n    ax.set_yticklabels(names, fontsize=12.5, color=INK)\n    top = float(values.max()) if len(values) else 1.0\n    pad = top * pad_frac\n    for yy, v, txt in zip(y, values, annot):\n        ax.text(v + pad, yy, txt, va=\"center\", ha=\"left\", fontsize=12,\n                color=INK_2)\n    ax.set_xlim(0, top * xpad)\n    ax.set_ylim(-0.75, len(names) - 0.25)\n    style(ax, \"x\")\n    ax.spines[\"left\"].set_visible(False)\n    return bars\n\n\ndef family_legend(ax, loc=\"lower right\", ncol=1):\n    \"\"\"Attach the fixed anatomical-family colour legend to an axes.\n\n    Args:\n        ax: Target axes.\n        loc: Matplotlib legend location.\n        ncol: Number of legend columns.\n\n    Returns:\n        The legend object.\n    \"\"\"\n    handles = [Line2D([0], [0], marker=\"s\", linestyle=\"none\", markersize=11,\n                      markerfacecolor=c, markeredgecolor=\"none\", label=k)\n               for k, c in FAMILY_COLORS.items()]\n    return ax.legend(handles=handles, loc=loc, ncol=ncol, handletextpad=0.5,\n                     labelcolor=INK_2, borderaxespad=0.6)\n\n\ndef show_table(df, caption=\"\"):\n    \"\"\"Render a small dataframe as the table view that accompanies a chart.\n\n    Args:\n        df: The dataframe to display.\n        caption: Optional caption printed beneath.\n    \"\"\"\n    from IPython.display import display\n    display(df)\n    if caption:\n        print(caption)\n\n\nprint(\"design system ready:\", len(FAMILY_COLORS), \"family hues,\",\n      len(BLUE_STEPS), \"sequential steps\")\nprint(\"text contrast on SURFACE (WCAG AA floor 4.5):\",\n      \", \".join(f\"{n} {contrast(c, SURFACE):.1f}\" for n, c in\n                [(\"INK\", INK), (\"INK_2\", INK_2), (\"MUTED\", MUTED),\n                 (\"ACCENT_TEXT\", ACCENT_TEXT)]))"},{"cell_type":"code","execution_count":null,"id":"7ac5d34b","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Load the five competition CSVs.\"\"\"\n\ntrain = pd.read_csv(ROOT / \"train.csv\")\ntrain_series = pd.read_csv(ROOT / \"train_series.csv\")\ntest = pd.read_csv(ROOT / \"test.csv\")\ntest_series = pd.read_csv(ROOT / \"test_series.csv\")\nsample_sub = pd.read_csv(ROOT / \"sample_submission.csv\")\n\nfor name, df in [(\"train\", train), (\"train_series\", train_series),\n                 (\"test\", test), (\"test_series\", test_series),\n                 (\"sample_submission\", sample_sub)]:\n    print(f\"{name:18s} {df.shape[0]:>7,} x {df.shape[1]:<3d}  {list(df.columns)[:4]}\")\n\n# The submission column order is fixed by sample_submission.csv. Check that the\n# label list used throughout this notebook matches it exactly.\nassert list(sample_sub.columns) == [\"StudyInstanceUID\"] + LABELS, \\\n    \"submission column order does not match the LABELS constant\"\nprint(\"\\nsubmission column order matches LABELS: OK\")"},{"cell_type":"markdown","id":"87ba5d65","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">1</span><span style=\"vertical-align:middle;\">The shape of the problem</span></h2></div>\n\nOne study is one prediction row of twelve probabilities. Underneath each study sit five or\nsix series, and underneath each series sit twenty to forty-five slices, with a tail that\nruns into the hundreds. The scored unit is therefore three levels above the unit the model\nactually sees, and something has to aggregate slices to series to study."},{"cell_type":"code","execution_count":null,"id":"b967123e","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Headline inventory as stat tiles. The number is the chart.\"\"\"\n\nn_studies = train[\"StudyInstanceUID\"].nunique()\nn_series = len(train_series)\nn_test_series = len(test_series)\n\ntiles = [\n    (f\"{n_studies:,}\", \"training studies\", \"one prediction row each\"),\n    (f\"{n_series:,}\", \"training series\", f\"{n_series / n_studies:.2f} per study\"),\n    (\"12\", \"findings per study\", \"each scored separately\"),\n    (\"macro AUC\", \"the metric\", \"a rare label costs as much as a common one\"),\n]\n\n# One axes, four cards drawn on it. These are stat tiles, not analysis panels:\n# no axis, no marks, nothing to compare along a shared scale. Drawing them as\n# rectangles on a single axes rather than as a 2 x 2 grid of subplots keeps the\n# notebook's one-or-two-panel rule literally true, and a 2 x 2 block of cards\n# reads better than four thin cards across a wide column.\nfig, ax = plt.subplots(figsize=(13.5, 5.4))\nax.axis(\"off\")\nax.set_xlim(0, 1)\nax.set_ylim(0, 1)\n# The gutters are asymmetric in axes fraction because the axes is about three\n# times wider than it is tall; both work out to roughly a fifth of an inch.\nCARD_W, GAP_X = 0.492, 0.016\nGAP_Y = 0.055\nCARD_H = (1.0 - GAP_Y) / 2.0\nfor k, (big, mid, small) in enumerate(tiles):\n    x0 = 0.0 if k % 2 == 0 else CARD_W + GAP_X\n    y0 = CARD_H + GAP_Y if k < 2 else 0.0\n    ax.add_patch(plt.Rectangle((x0, y0), CARD_W, CARD_H, facecolor=SURFACE,\n                               edgecolor=GRID, linewidth=1.2, zorder=0))\n    ax.add_patch(plt.Rectangle((x0, y0), 0.011, CARD_H, facecolor=PRIMARY,\n                               edgecolor=\"none\", zorder=1))\n    size = 38 if len(big) <= 6 else 25\n    tx = x0 + 0.030\n    ax.text(tx, y0 + CARD_H * 0.70, big, fontsize=size, fontweight=\"bold\",\n            color=INK, va=\"center\", ha=\"left\")\n    ax.text(tx, y0 + CARD_H * 0.37, mid, fontsize=14, color=INK_2,\n            va=\"center\", ha=\"left\")\n    ax.text(tx, y0 + CARD_H * 0.17, small, fontsize=12, color=MUTED,\n            va=\"center\", ha=\"left\")\n\nfinish(fig, \"4,407 studies, 24,371 series, 12 findings, one macro-averaged AUC\",\n       detail=\"The scored unit is the study. The unit a network sees is the slice, \"\n              \"three levels down.\",\n       caption=f\"Counts read from train.csv and train_series.csv at run time. \"\n               f\"test_series.csv ships {n_test_series} placeholder rows.\")"},{"cell_type":"markdown","id":"61d7ef8d","metadata":{},"source":"**Finding.** Macro averaging is the design decision that shapes effort allocation. A label\nwith nine positives in the gold set is worth exactly as much of the final score as\nEffusion, which has thirty-five. A model that is near-random on one finding gives up a\ntwelfth of the achievable margin above 0.5, which is 0.042 of the final score, no matter\nhow good it is on the other eleven. Rare findings deserve disproportionate attention, not\nless."},{"cell_type":"markdown","id":"c20188a8","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">2</span><span style=\"vertical-align:middle;\">The label problem</span></h2></div>\n\nThis is the section that decides how the competition is won or lost."},{"cell_type":"code","execution_count":null,"id":"2dbfe37d","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"How many studies actually carry labels?\"\"\"\n\nlabel_block = train[LABELS]\nn_labelled_per_row = label_block.notna().sum(axis=1)\ngold = train.loc[n_labelled_per_row == 12].reset_index(drop=True)\nunlabelled = train.loc[n_labelled_per_row == 0]\n\nprint(\"label completeness by row\")\nprint(n_labelled_per_row.value_counts().rename(\"studies\").to_string())\nprint()\nprint(f\"fully labelled studies : {len(gold):,}  ({100 * len(gold) / len(train):.2f}%)\")\nprint(f\"unlabelled studies     : {len(unlabelled):,}\")\nprint(f\"partially labelled     : {len(train) - len(gold) - len(unlabelled):,}\")\nprint()\nprint(\"distinct values in the labelled block:\",\n      sorted(pd.unique(gold[LABELS].to_numpy().ravel())))\n\n# Every label column is non-null on exactly the same rows?\nnn_sets = {c: frozenset(train.index[train[c].notna()]) for c in LABELS}\nsame = len(set(nn_sets.values())) == 1\nprint(\"all twelve label columns share one non-null row set:\", same)"},{"cell_type":"code","execution_count":null,"id":"bf706c2a","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Unit chart: one square per training study, the labelled ones highlighted.\"\"\"\n\nn_total = len(train)\nncols = 147\nnrows = int(np.ceil(n_total / ncols))\ngrid = np.full(nrows * ncols, np.nan)\ngrid[:len(gold)] = 1.0          # labelled, grouped for legibility\ngrid[len(gold):n_total] = 0.0   # unlabelled\ngrid = grid.reshape(nrows, ncols)\n\nfig, ax = plt.subplots(figsize=(14.0, 6.0))\ncmap = mpl.colors.ListedColormap([GRID, ACCENT])\nax.pcolormesh(np.arange(ncols + 1), np.arange(nrows + 1)[::-1], grid,\n              cmap=cmap, vmin=0, vmax=1, edgecolors=SURFACE, linewidth=0.45)\nax.set_xlim(-1, ncols + 1)\nax.set_ylim(-5.0, nrows + 4.2)\nax.set_aspect(\"equal\")\nax.axis(\"off\")\n\nax.annotate(f\"{len(gold)} studies carry all twelve labels\",\n            xy=(len(gold) / 2.0, nrows - 0.1),\n            xytext=(len(gold) / 2.0, nrows + 2.6),\n            fontsize=14.5, color=ACCENT_TEXT, fontweight=\"bold\", va=\"bottom\",\n            ha=\"center\",\n            arrowprops=dict(arrowstyle=\"-\", color=ACCENT_TEXT, linewidth=1.4))\nax.text(ncols, -1.4, f\"{len(unlabelled):,} studies carry a report and nothing else\",\n        fontsize=13.5, color=INK_2, va=\"top\", ha=\"right\")\n\nfinish(fig, \"Only 1.3% of training studies carry labels\",\n       detail=f\"{len(gold)} of {n_total:,} studies have all twelve findings filled in. \"\n              f\"The other {len(unlabelled):,} ship a free-text radiology report and nothing else.\",\n       caption=\"One square is one study. The labelled studies are grouped at the top left \"\n               \"for legibility; in the file they are scattered.\")"},{"cell_type":"markdown","id":"6e1f45d5","metadata":{},"source":"**Finding.** The imaging model is downstream of a label pipeline that does not exist yet.\nTraining a network on 58 studies is not a serious plan for a twelve-label problem; training\nit on 4,407 studies requires turning 4,349 radiology reports into 52,188 binary decisions\nfirst. Label quality caps model quality, and the ceiling is set before a single convolution\nruns. Budget accordingly: the extraction step deserves more effort than the backbone\nchoice."},{"cell_type":"code","execution_count":null,"id":"818a853f","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Per-label positive counts among the 58 gold studies.\"\"\"\n\npos = gold[LABELS].sum().astype(int)\norder = pos.sort_values(ascending=False).index.tolist()\nvals = pos[order].to_numpy()\ncols = [LABEL_COLORS[l] for l in order]\n\nfig, ax = plt.subplots(figsize=(14.0, 6.8))\nhbar(ax, order, vals, cols,\n     annot=[f\"{v}   {100 * v / len(gold):.0f}%\" for v in vals], xpad=1.30)\nax.set_xlabel(f\"positive studies (of {len(gold)})\")\nfamily_legend(ax, loc=\"lower right\")\n\nfinish(fig, \"Every finding has between 9 and 35 positives in the gold set\",\n       detail=\"MCL is the scarcest at 9. Effusion is the most common at 35. \"\n              \"Both count equally toward the macro average.\",\n       caption=\"Colour marks the anatomical family, not the individual finding; \"\n               \"twelve distinct hues would fail colourblind separation. \"\n               \"Percentages are of the 58 gold studies.\")\n\nshow_table(pd.DataFrame({\n    \"positives\": pos[order],\n    \"negatives\": len(gold) - pos[order],\n    \"prevalence\": (pos[order] / len(gold)).round(3),\n    \"family\": [LABEL_FAMILY[l] for l in order],\n}))"},{"cell_type":"markdown","id":"53360596","metadata":{},"source":"**Finding.** MCL carries nine positives. A single mislabelled study moves its AUC by\nseveral points. Any per-label metric computed here is a coin flip dressed as a decimal, and\nthe two scarcest labels, MCL and Lateral OA, together account for a sixth of the macro\nscore."},{"cell_type":"code","execution_count":null,"id":"f13349a4","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Is the gold set a random sample of studies? Findings per study, and co-occurrence.\"\"\"\n\nper_study = gold[LABELS].sum(axis=1)\n\nfig = plt.figure(figsize=(13.5, 11.6))\ngs = fig.add_gridspec(2, 1, height_ratios=[1.0, 1.30], hspace=0.30)\n\nax0 = fig.add_subplot(gs[0, 0])\ncounts = per_study.value_counts().sort_index()\nxs = np.arange(0, 13)\nys = np.array([counts.get(float(k), 0) for k in xs])\nax0.bar(xs, ys, width=0.72, color=PRIMARY, zorder=3)\ntop_y = ys.max()\nax0.set_ylim(0, top_y * 1.40)\nax0.plot([0], [0], marker=\"o\", markersize=9, color=ACCENT, zorder=5,\n         markeredgecolor=SURFACE, markeredgewidth=2)\n# A vertical leader, drawn explicitly. An annotate() arrow to a left-aligned\n# label lands on the far edge of the text box and draws a long diagonal across\n# the bars, which reads as a data series. The column above x=0 is empty, so a\n# straight rule there crosses nothing.\nax0.plot([0.0, 0.0], [top_y * 0.10, top_y * 1.15], color=ACCENT, linewidth=1.3,\n         solid_capstyle=\"butt\", zorder=4)\nax0.text(0.12, top_y * 1.25, \"the bin at zero is empty\", fontsize=12.5,\n         color=ACCENT_TEXT, fontweight=\"bold\", va=\"center\", ha=\"left\")\nfor x, y in zip(xs, ys):\n    if y:\n        ax0.text(x, y + top_y * 0.025, str(int(y)), ha=\"center\", va=\"bottom\",\n                 fontsize=12, color=INK_2)\nax0.axvline(per_study.mean(), color=ACCENT, linewidth=2.0, zorder=4)\nax0.text(per_study.mean() + 0.2, top_y * 1.05, f\"mean {per_study.mean():.2f}\",\n         color=ACCENT_TEXT, fontsize=12.5, fontweight=\"bold\", va=\"center\")\nax0.set_xticks(xs)\nax0.set_xlabel(\"positive findings in the study\")\nax0.set_ylabel(\"gold studies\")\nax0.set_title(\"Not one study is negative on all twelve\", loc=\"left\",\n              fontsize=13.5, color=INK, pad=10)\nstyle(ax0, \"y\")\n\nax1 = fig.add_subplot(gs[1, 0])\nY = gold[LABELS].to_numpy()\nco = Y.T @ Y\n# aspect=\"auto\" rather than the imshow default: stacked under a full-width bar\n# chart, a square 12 x 12 matrix would sit in a narrow column with empty page\n# either side. Wide cells cost nothing here because every cell is annotated.\ntop_co = float(co.max()) or 1.0\nim = ax1.imshow(co, cmap=SEQ_SAFE, vmin=0, vmax=top_co, aspect=\"auto\")\nax1.set_xticks(range(12))\nax1.set_yticks(range(12))\nax1.set_xticklabels(LABELS, rotation=45, ha=\"right\", fontsize=11, color=INK_2)\nax1.set_yticklabels(LABELS, fontsize=11, color=INK_2)\nfor i in range(12):\n    for j in range(12):\n        ax1.text(j, i, int(co[i, j]), ha=\"center\", va=\"center\", fontsize=10.5,\n                 color=ink_on(SEQ_SAFE(co[i, j] / top_co)))\nfor s in ax1.spines.values():\n    s.set_visible(False)\nax1.tick_params(length=0)\nax1.set_title(\"Co-occurrence counts; the diagonal is the label's own prevalence\",\n              loc=\"left\", fontsize=13.5, color=INK, pad=10)\n\nfinish(fig, \"The gold 58 are positive-enriched, not a random sample of studies\",\n       detail=f\"Every one carries at least one finding, and the mean study carries \"\n              f\"{per_study.mean():.2f} of 12.\",\n       headroom=0.38, footroom=1.35,\n       caption=\"Top: positive findings per gold study. Bottom: pairwise co-occurrence \"\n               \"among the 58; every cell is annotated, so the chart is its own table view.\")"},{"cell_type":"markdown","id":"65c57077","metadata":{},"source":"**Finding.** The bin at zero findings is empty. Selection for the gold set was conditional\non abnormality, so prevalence in these 58 studies tells you nothing about prevalence in the\n4,349 unlabelled ones, and nothing about the test set. Two consequences follow immediately.\nDo not calibrate probabilities to gold-set prevalence. Do not use the gold set to estimate\nhow often a finding occurs, only to check whether an extracted label agrees with a\nradiologist when the two are compared on the same study.\n\nThe competition host issues a matching warning: prevalence is not guaranteed to be the same\nacross training, public leaderboard, and final evaluation. AUC is insensitive to prevalence\nin expectation, so this is mostly a warning about variance and about tuning thresholds."},{"cell_type":"code","execution_count":null,"id":"3c42cd68","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"How precisely can 58 studies measure an AUC? Two independent answers.\"\"\"\n\nfrom scipy.stats import norm, rankdata\n\n\ndef hanley_se(auc, n_pos, n_neg):\n    \"\"\"Hanley-McNeil standard error of an empirical ROC AUC.\n\n    Args:\n        auc: Assumed true area under the ROC curve.\n        n_pos: Number of positive cases.\n        n_neg: Number of negative cases.\n\n    Returns:\n        The standard error of the empirical AUC estimate.\n    \"\"\"\n    q1 = auc / (2.0 - auc)\n    q2 = 2.0 * auc * auc / (1.0 + auc)\n    num = (auc * (1 - auc) + (n_pos - 1) * (q1 - auc ** 2)\n           + (n_neg - 1) * (q2 - auc ** 2))\n    return float(np.sqrt(num / (n_pos * n_neg)))\n\n\ndef fast_auc(y, s):\n    \"\"\"ROC AUC via the Mann-Whitney rank statistic, with tie correction.\n\n    Args:\n        y: Binary ground truth, shape ``(n,)``.\n        s: Scores, shape ``(n,)``.\n\n    Returns:\n        The AUC, or ``float(\"nan\")`` when one class is absent.\n    \"\"\"\n    y = np.asarray(y)\n    n1 = int(y.sum())\n    n0 = len(y) - n1\n    if n1 == 0 or n0 == 0:\n        return float(\"nan\")\n    r = rankdata(s)\n    return float((r[y == 1].sum() - n1 * (n1 + 1) / 2.0) / (n1 * n0))\n\n\nTRUE_AUC = 0.80  # the reference operating point the intervals are computed at\n\nrows = []\nfor lab in LABELS:\n    p = int(pos[lab])\n    n = len(gold) - p\n    se = hanley_se(TRUE_AUC, p, n)\n    rows.append({\"label\": lab, \"pos\": p, \"neg\": n, \"se\": se, \"half_ci\": 1.96 * se})\nci = pd.DataFrame(rows).set_index(\"label\").loc[LABELS]\n\n# Bootstrap the macro average over studies, so label correlation is preserved.\nmu = np.sqrt(2.0) * norm.ppf(TRUE_AUC)\nboot_rng = np.random.default_rng(SEED)\nscores = boot_rng.normal(loc=Y * mu, scale=1.0)\nidx = np.arange(len(gold))\nmacro = []\nfor _ in range(N_BOOTSTRAP):\n    b = boot_rng.choice(idx, size=len(gold), replace=True)\n    a = [fast_auc(Y[b, j], scores[b, j]) for j in range(12)]\n    if not np.any(np.isnan(a)):\n        macro.append(float(np.mean(a)))\nmacro = np.asarray(macro)\nlo, hi = np.percentile(macro, [2.5, 97.5])\n\nprint(f\"assumed true AUC per label      : {TRUE_AUC:.2f}\")\nprint(f\"mean per-label standard error   : {ci['se'].mean():.4f}\")\nprint(f\"widest per-label 95% interval   : +/- {ci['half_ci'].max():.3f} ({ci['half_ci'].idxmax()})\")\nprint(f\"bootstrap macro AUC 95% interval: [{lo:.4f}, {hi:.4f}]  width {hi - lo:.4f}\")\nprint(f\"bootstrap resamples used        : {len(macro):,} of {N_BOOTSTRAP:,}\")"},{"cell_type":"code","execution_count":null,"id":"315ad9d3","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Chart the noise floor: per-label intervals, and the macro-average interval.\"\"\"\n\nfig = plt.figure(figsize=(13.5, 11.2))\ngs = fig.add_gridspec(2, 1, height_ratios=[1.35, 1.0], hspace=0.32)\n\nax0 = fig.add_subplot(gs[0, 0])\no = ci[\"half_ci\"].sort_values(ascending=False).index.tolist()\nypos = np.arange(len(o))[::-1]\nfor yy, lab in zip(ypos, o):\n    h = ci.loc[lab, \"half_ci\"]\n    c = LABEL_COLORS[lab]\n    ax0.plot([TRUE_AUC - h, TRUE_AUC + h], [yy, yy], color=c, linewidth=3.0,\n             solid_capstyle=\"round\", zorder=3)\n    ax0.plot([TRUE_AUC], [yy], marker=\"o\", color=c, markersize=9,\n             markeredgecolor=SURFACE, markeredgewidth=2, zorder=4)\n    ax0.text(TRUE_AUC + h + 0.012, yy, f\"±{h:.3f}\", va=\"center\", ha=\"left\",\n             fontsize=12, color=INK_2)\nax0.axvspan(TRUE_AUC - 0.005, TRUE_AUC + 0.005, color=INK, alpha=0.10, zorder=1)\nax0.set_yticks(ypos)\nax0.set_yticklabels(o, fontsize=12.5, color=INK)\n# 1.06, not 1.02: the widest interval reaches 0.982 and its ±0.182 label has to\n# sit to the right of it without running off the panel.\nax0.set_xlim(0.55, 1.06)\nax0.set_ylim(-2.4, len(o) - 0.4)\nax0.set_xlabel(f\"AUC measured on {len(gold)} studies, if the true AUC were {TRUE_AUC:.2f}\")\nstyle(ax0, \"x\")\nax0.spines[\"left\"].set_visible(False)\nax0.set_title(\"Per-label 95% intervals (Hanley-McNeil)\", loc=\"left\",\n              fontsize=13.5, color=INK, pad=10)\nax0.text(TRUE_AUC + 0.014, -0.80, \"the shaded band is 0.01 wide, roughly the\\n\"\n         \"margin that separates leaderboard positions\",\n         fontsize=11.25, color=MUTED, va=\"top\")\n\nax1 = fig.add_subplot(gs[1, 0])\nax1.hist(macro, bins=44, color=PRIMARY, edgecolor=SURFACE, linewidth=0.5, zorder=3)\nymax = ax1.get_ylim()[1]\nax1.set_ylim(0, ymax * 1.30)\nax1.axvline(lo, color=ACCENT, linewidth=2.0, zorder=4)\nax1.axvline(hi, color=ACCENT, linewidth=2.0, zorder=4)\nax1.annotate(\"\", xy=(lo, ymax * 1.09), xytext=(hi, ymax * 1.09),\n             arrowprops=dict(arrowstyle=\"<->\", color=ACCENT, linewidth=1.6))\nax1.text((lo + hi) / 2, ymax * 1.14, f\"95% interval spans {hi - lo:.3f}\",\n         ha=\"center\", va=\"bottom\", fontsize=13, color=ACCENT_TEXT,\n         fontweight=\"bold\")\nax1.set_xlabel(\"macro AUC measured on the gold 58\")\nax1.set_ylabel(\"bootstrap resamples\")\nstyle(ax1, \"y\")\nax1.set_title(\"Macro average, bootstrapped over studies\", loc=\"left\",\n              fontsize=13.5, color=INK, pad=10)\n\nfinish(fig, f\"The gold 58 resolve macro AUC only to about ±{(hi - lo) / 2:.02f}, \"\n            f\"so it cannot rank models\",\n       detail=f\"Per-label intervals run from ±{ci['half_ci'].min():.3f} to \"\n              f\"±{ci['half_ci'].max():.3f}.\",\n       headroom=0.38,\n       caption=f\"Scores simulated bi-normal at a true AUC of {TRUE_AUC:.2f}, seed {SEED}; \"\n               f\"{len(macro):,} bootstrap resamples over studies, so label co-occurrence \"\n               f\"is preserved. The shaded band in the top panel is 0.01 wide.\")"},{"cell_type":"markdown","id":"dfcccaa5","metadata":{},"source":"**Finding.** A 95% interval of roughly 0.08 on the macro average is the single most\nimportant number in this notebook. Competitions of this kind are decided on differences of\n0.005 to 0.02. Choosing a backbone, a slice-aggregation scheme, or a learning rate by its\nscore on the gold 58 is choosing by coin flip, and it will feel like signal the whole time.\n\nUse the gold 58 for the one thing 58 studies can support: auditing the report-to-label\nextraction. Per label, on these 58 studies, does the extractor agree with the radiologist?\nThat is a per-label agreement rate with an honest confidence interval, and it is the number\nthat tells you whether the 4,349 derived labels are worth training on.\n\nModel ranking has to come from a much larger validation set built out of report-derived\nlabels, with folds grouped by patient. Section 6 shows why the grouping key is `PatientID`\nand not `StudyInstanceUID`."},{"cell_type":"markdown","id":"6913e4e0","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">3</span><span style=\"vertical-align:middle;\">The reports</span></h2></div>\n\n4,407 reports, one per study, all present. They are the only path to labels for 98.7% of\nthe training set, so their structure is not a curiosity, it is the input to the most\nimportant component in the pipeline."},{"cell_type":"code","execution_count":null,"id":"3aeb82b1","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Report length distribution.\"\"\"\n\nreports = train[\"Report\"].fillna(\"\")\nlengths = reports.str.len()\nwords = reports.str.split().str.len()\n\nprint(f\"reports present : {int(train['Report'].notna().sum()):,} of {len(train):,}\")\nprint(f\"empty strings   : {int((lengths == 0).sum()):,}\")\nq = lengths.quantile([0.01, 0.05, 0.25, 0.50, 0.75, 0.95, 0.99])\n\nfig, ax = plt.subplots(figsize=(14.0, 5.8))\nax.hist(lengths, bins=90, range=(0, 4000), color=PRIMARY, edgecolor=SURFACE,\n        linewidth=0.4, zorder=3)\nmed = lengths.median()\nax.axvline(med, color=ACCENT, linewidth=2.0, zorder=4)\nax.text(med + 40, ax.get_ylim()[1] * 0.92, f\"median {med:,.0f} characters\",\n        color=ACCENT_TEXT, fontsize=13, fontweight=\"bold\", va=\"center\")\np95 = lengths.quantile(0.95)\nax.axvline(p95, color=INK_2, linewidth=1.4, linestyle=\"-\", zorder=4)\nax.text(p95 + 40, ax.get_ylim()[1] * 0.72, f\"p95 {p95:,.0f}\", color=INK_2,\n        fontsize=12.5, va=\"center\")\nax.set_xlabel(\"report length (characters)\")\nax.set_ylabel(\"reports\")\nstyle(ax, \"y\")\n\nfinish(fig, \"Reports are short: half are under 1,000 characters\",\n       detail=f\"Range {lengths.min():,.0f} to {lengths.max():,.0f} characters, \"\n              f\"median {med:,.0f}, p95 {p95:,.0f}. Median {words.median():,.0f} words.\",\n       caption=\"Histogram truncated at 4,000 characters for legibility; \"\n               f\"{int((lengths > 4000).sum())} reports are longer.\")\n\nshow_table(q.rename(\"characters\").to_frame().assign(\n    words=words.quantile([0.01, 0.05, 0.25, 0.50, 0.75, 0.95, 0.99]).values).round(0))"},{"cell_type":"markdown","id":"338f3aa9","metadata":{},"source":"**Finding.** These fit comfortably in any context window, so the extraction step is not\nconstrained by document length. A median report of about 1,000 characters costs very little\nto process one at a time, and the entire corpus is a few million characters. Length is not\nthe obstacle. Language is."},{"cell_type":"code","execution_count":null,"id":"1684f07a","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"A language fingerprint: unambiguous script blocks, then a stopword vote.\"\"\"\n\nGREEK_RE = re.compile(r\"[Ͱ-Ͽἀ-῿]\")\nCYRILLIC_RE = re.compile(r\"[Ѐ-ӿ]\")\nCJK_RE = re.compile(r\"[぀-ヿ一-鿿가-힯]\")\nARABIC_RE = re.compile(r\"[؀-ۿ]\")\nTURKISH_DIAC_RE = re.compile(r\"[şŞğĞıİ]\")\nWORD_RE = re.compile(r\"[^\\W\\d_]+\", re.UNICODE)\n\nSTOPWORDS = {\n    \"English\": [\"the\", \"and\", \"with\", \"there\", \"normal\", \"signal\", \"joint\", \"tear\", \"no\"],\n    \"Spanish\": [\"de\", \"la\", \"el\", \"con\", \"sin\", \"rodilla\", \"menisco\", \"los\", \"una\"],\n    \"Turkish\": [\"ve\", \"ile\", \"yok\", \"diz\", \"normaldir\", \"izlenmektedir\", \"bulgular\"],\n    \"German\": [\"der\", \"die\", \"das\", \"und\", \"mit\", \"kein\", \"keine\", \"nicht\", \"des\"],\n    \"Dutch\": [\"het\", \"een\", \"van\", \"met\", \"geen\", \"knie\", \"niet\", \"zijn\"],\n    \"French\": [\"le\", \"la\", \"les\", \"avec\", \"sans\", \"genou\", \"est\", \"des\", \"une\"],\n    \"Italian\": [\"del\", \"della\", \"con\", \"senza\", \"ginocchio\", \"non\", \"menisco\"],\n    \"Portuguese\": [\"do\", \"da\", \"com\", \"sem\", \"joelho\", \"nao\", \"menisco\"],\n}\n\n\ndef fingerprint(text):\n    \"\"\"Guess a report's language from its script, then a stopword vote.\n\n    Script detection is exact; the stopword vote is a crude heuristic and is\n    labelled as such wherever its output is shown.\n\n    Args:\n        text: Raw report text.\n\n    Returns:\n        A language or script name, or ``\"Unresolved\"`` when no vote wins clearly.\n    \"\"\"\n    if GREEK_RE.search(text):\n        return \"Greek\"\n    if CYRILLIC_RE.search(text):\n        return \"Cyrillic script\"\n    if CJK_RE.search(text):\n        return \"CJK script\"\n    if ARABIC_RE.search(text):\n        return \"Arabic script\"\n    toks = collections.Counter(w.lower() for w in WORD_RE.findall(text))\n    if not toks:\n        return \"Unresolved\"\n    score = {k: sum(toks[w] for w in ws) for k, ws in STOPWORDS.items()}\n    if TURKISH_DIAC_RE.search(text):\n        score[\"Turkish\"] += 5\n    ranked = sorted(score.items(), key=lambda kv: -kv[1])\n    if ranked[0][1] < 3 or ranked[0][1] == ranked[1][1]:\n        return \"Unresolved\"\n    return ranked[0][0]\n\n\nlang = reports.map(fingerprint)\nlang_counts = lang.value_counts()\nprint(lang_counts.to_string())\nprint(f\"\\nnon-English share: {100 * (1 - lang_counts.get('English', 0) / len(lang)):.1f}%\")"},{"cell_type":"code","execution_count":null,"id":"4ab2ee66","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Chart the language mix, emphasising the one thing that matters.\"\"\"\n\nlc = lang_counts.sort_values(ascending=False)\nnames = lc.index.tolist()\nvals = lc.to_numpy()\n\n# Reading the members of these two buckets shows the vote is wrong for both:\n# the Italian bucket holds Spanish, the Portuguese bucket holds Croatian. Mark\n# them as unreliable in the chart rather than leaving the reader to trust them.\nSUSPECT = {\"Italian\", \"Portuguese\", \"Unresolved\"}\ncols = [ACCENT if n == \"English\" else NEUTRAL if n in SUSPECT else PRIMARY\n        for n in names]\n\nfig, ax = plt.subplots(figsize=(14.0, 6.2))\nhbar(ax, names, vals, cols,\n     annot=[f\"{v:,}   {100 * v / len(lang):.0f}%\" for v in vals], xpad=1.30)\nax.set_xlabel(\"reports\")\nhandles = [Line2D([0], [0], marker=\"s\", linestyle=\"none\", markersize=11, label=t,\n                  markerfacecolor=c, markeredgecolor=\"none\")\n           for t, c in [(\"English\", ACCENT), (\"another identified language\", PRIMARY),\n                        (\"vote unreliable or absent\", NEUTRAL)]]\nax.legend(handles=handles, loc=\"lower right\", labelcolor=INK_2, borderaxespad=0.8)\n\nfinish(fig, f\"English is a plurality, not a majority: \"\n            f\"{100 * (1 - lc.get('English', 0) / len(lang)):.0f}% of reports are not in English\",\n       detail=\"Greek and Cyrillic script are detected exactly. The Latin-script split \"\n              \"comes from a stopword vote and is approximate.\",\n       caption=\"This fingerprint is deliberately crude and its bucket sizes shift with the \"\n               \"token list. What survives any token list is the claim that matters: at \"\n               \"least eight languages are present and no single one dominates.\")"},{"cell_type":"markdown","id":"830ead5d","metadata":{},"source":"**Finding.** A label extractor tuned on English clears roughly 40% of the corpus and\nsilently fails on the rest. Failure will not look like an error, it will look like a\nnegative label, because a keyword that never matches produces a zero. That is the worst\npossible failure mode: quiet, systematic, and correlated with imaging site.\n\nThe fingerprint above is itself unreliable, which is part of the finding. Its Italian and\nPortuguese buckets are wrong on inspection: the Italian bucket catches Spanish text, and\nthe Portuguese bucket catches Croatian. Language identification on short clinical text with\nheavy Latin abbreviation is not a solved sub-problem you can bolt on. Design the extraction\nso it does not need to know the language."},{"cell_type":"code","execution_count":null,"id":"9f9c32d1","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Verbatim report openings across languages. Truncated, never edited.\"\"\"\n\n\ndef show_examples(picker, n_langs, width=300):\n    \"\"\"Print one truncated verbatim report per requested language bucket.\n\n    Args:\n        picker: Series of language labels aligned with ``reports``.\n        n_langs: Ordered list of bucket names to show.\n        width: Characters kept from the start of each report.\n    \"\"\"\n    for name in n_langs:\n        sub = reports[picker == name]\n        if not len(sub):\n            print(f\"--- {name}: no report matched ---\\n\")\n            continue\n        txt = sub.iloc[0][:width].replace(\"\\n\", \" \").strip()\n        print(f\"--- {name}  (n={len(sub):,}) \" + \"-\" * max(0, 52 - len(name)))\n        print(txt + (\" ...\" if len(sub.iloc[0]) > width else \"\"))\n        print()\n\n\nshow_examples(lang, [\"English\", \"Spanish\", \"Turkish\", \"Greek\", \"German\",\n                     \"Dutch\", \"French\", \"Cyrillic script\", \"Unresolved\"])"},{"cell_type":"markdown","id":"2a787274","metadata":{},"source":"**Finding.** These are not translations of one template. The Turkish report is written as a\nprotocol block followed by a run of `normaldir` assertions. The Greek report opens with a\n`ΤΕΧΝΙΚΗ` section and states absence with `Χωρίς`. The Dutch reports carry a Belgian\nhospital header with `Klinische inlichtingen` and `Bevindingen` sections. The French one\nuses a section-header style where the heading names the finding and the body says whether\nit is there. Section structure, negation form, and terminology all vary, and there is no\nshared skeleton to parse against."},{"cell_type":"code","execution_count":null,"id":"f2f7ac87","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Worked example: why keyword extraction fails. Fracture as the test case.\"\"\"\n\n# Both Greek mu forms are needed here; see the homoglyph check at the end of\n# this cell for why.\nFRACTURE_TOKEN = re.compile(\n    r\"fractur|fraktur|fraktür|kırık|kirik|κάταγ[μµ]|καταγ[μµ]|перелом|прелом\",\n    re.IGNORECASE)\n\n# Deliberately narrow. Earlier drafts of this cell included \"izlenme\" as a\n# Turkish negator, which is wrong: \"izlenmektedir\" means the finding IS seen and\n# \"izlenmemektedir\" means it is not. The loose list overstated negation.\nNEGATORS = [\n    \"aucune\", \"aucun\", \"pas de\",\n    \"no \", \"not \", \"none\", \"without\", \"absence\", \"negative for\", \"unremarkable\",\n    \"kein\", \"keine\",\n    \"geen\",\n    \"sin \", \"no se observ\", \"no hay\",\n    \"izlenmemekte\", \"izlenmemiş\", \"saptanmamış\", \"saptanmadı\", \"görülmemekte\",\n    \"yoktur\", \" yok\", \"mevcut değil\",\n    \"χωρίς\", \"δεν \", \"ουδεμία\",\n    \"нет\", \"без\", \"не выявлен\", \"не определ\",\n]\nPRE, POST = 40, 60   # negator search window, characters either side of the token\n\nmentions = reports.str.contains(FRACTURE_TOKEN, regex=True)\nn_mentions = int(mentions.sum())\n\nnegated = 0\nfor txt in reports[mentions]:\n    low = txt.lower()\n    for m in FRACTURE_TOKEN.finditer(low):\n        window = low[max(0, m.start() - PRE):m.end() + POST]\n        if any(x in window for x in NEGATORS):\n            negated += 1\n            break\n\nprint(f\"reports containing a fracture token        : {n_mentions:,}\")\nprint(f\"of those, a negator sits within the window : {negated:,} \"\n      f\"({100 * negated / max(n_mentions, 1):.0f}%)\")\nprint(f\"gold studies actually positive for Fracture: {int(pos['Fracture'])} of {len(gold)} \"\n      f\"({100 * pos['Fracture'] / len(gold):.0f}%)\")\nprint()\n\n# Every entry below is a genuine negation, verified by reading it.\nPROBES = [\n    (\"French\", \"fracture\", r\"Fractures?\\s*:\\s*\\n?\\s*Aucune\"),\n    (\"English\", \"fracture\", r\"[Nn]o (acute |displaced |evidence of )?fractur\"),\n    (\"German\", \"fracture\", r\"[Kk]eine\\s+\\w*[Ff]raktur\"),\n    (\"Spanish\", \"fracture\", r\"[Ss]in\\s+\\w*fractur|[Nn]o\\s+se\\s+observan?\\s+\\w*fractur\"),\n    (\"Turkish\", \"fracture\", r\"kırık[^.\\n]{0,30}yok|fraktür[^.\\n]{0,30}yok\"),\n    (\"Greek\", \"effusion\", r\"[Χχ]ωρίς\\s+ενδαρθρική[^.\\n]{0,30}\"),\n    (\"Bulgarian\", \"meniscal tear\", r\"без\\s+данни\\s+за[^.\\n]{0,30}\"),\n    (\"Dutch\", \"ligament tear\", r\"[Kk]ruisbanden[^.\\n]{0,30}intact\"),\n]\nprint(\"Verbatim negations. A keyword extractor calls every one of these POSITIVE.\\n\")\nfor name, what, pat in PROBES:\n    rx = re.compile(pat)\n    hit = None\n    for txt in reports:\n        m = rx.search(txt)\n        if m:\n            s = max(0, m.start() - 30)\n            hit = txt[s:m.end() + 55].replace(\"\\n\", \" ⏎ \").strip()\n            break\n    body = f\"...{hit}...\" if hit else \"(no match for this probe)\"\n    print(f\"  [{name:10s} {what:14s}] {body}\")\n\n# A second trap, entirely separate from negation.\ngreek_reports = reports[reports.str.contains(GREEK_RE, regex=True)]\nmicro_sign = int(greek_reports.str.contains(\"µ\").sum())   # MICRO SIGN\ngreek_mu = int(greek_reports.str.contains(\"μ\").sum())     # GREEK SMALL MU\nprint(f\"\\nHomoglyph check across the {len(greek_reports):,} Greek reports:\")\nprint(f\"  contain MICRO SIGN U+00B5 (µ)        : {micro_sign:,} \"\n      f\"({100 * micro_sign / max(len(greek_reports), 1):.0f}%)\")\nprint(f\"  contain GREEK SMALL LETTER MU U+03BC : {greek_mu:,}\")\nprint(\"  A pattern written with the correct Greek mu misses most Greek reports.\")"},{"cell_type":"code","execution_count":null,"id":"5610ee6d","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Chart the gap between a token match and a true positive. One denominator.\"\"\"\n\nsurvives = n_mentions - negated\n\nfig, ax = plt.subplots(figsize=(14.0, 3.9))\nsegments = [(negated, \"a negator sits within the window\", ACCENT),\n            (survives, \"no negator nearby\", NEUTRAL)]\nleft = 0.0\nfor value, name, colour in segments:\n    ax.barh([0], [value], left=left, height=0.46, color=colour, zorder=3,\n            edgecolor=SURFACE, linewidth=2.4)\n    ax.text(left + value / 2.0, 0, f\"{value:,}\\n{100 * value / n_mentions:.0f}%\",\n            ha=\"center\", va=\"center\", fontsize=15, fontweight=\"bold\",\n            color=ink_on(colour))\n    left += value\nax.set_yticks([])\nax.set_ylim(-0.75, 0.55)\nax.set_xlim(0, n_mentions)\nax.set_xlabel(f\"reports containing a fracture token (n = {n_mentions:,})\")\nstyle(ax, \"none\")\nfor side in (\"left\", \"bottom\"):\n    ax.spines[side].set_visible(False)\nhandles = [Line2D([0], [0], marker=\"s\", linestyle=\"none\", markersize=11, label=n,\n                  markerfacecolor=c, markeredgecolor=\"none\")\n           for _, n, c in segments]\nax.legend(handles=handles, loc=\"lower left\", ncol=2, labelcolor=INK_2,\n          borderaxespad=0.2)\n\nfinish(fig, f\"{100 * negated / max(n_mentions, 1):.0f}% of the reports that mention a \"\n            f\"fracture negate it\",\n       detail=f\"A keyword extractor labels all {n_mentions:,} of these reports positive. \"\n              f\"At most {survives:,} could be.\",\n       caption=f\"Negator list covering eight languages, searched over a -{PRE}/+{POST} \"\n               f\"character window around each token; a lower bound on negation, not an \"\n               f\"exact count. For scale, {int(pos['Fracture'])} of the {len(gold)} \"\n               f\"positive-enriched gold studies are Fracture-positive.\")"},{"cell_type":"markdown","id":"055a2f95","metadata":{},"source":"**Finding.** The French example is the cleanest illustration. The report reads\n`Fractures :` on one line and `Aucune.` on the next. The word *fracture* is present, the\nfinding is absent, and the two are separated by a line break that any bag-of-words method\nthrows away. English has `no fracture or bone contusion`, where one negation scopes two\nfindings at once. Turkish puts `yok` after the noun. Greek uses `Χωρίς`. Bulgarian uses\n`без данни за`.\n\nNegation is not an edge case here, it is the majority case: about three quarters of the\nreports that mention a fracture go on to negate it.\n\nTwo details are worth dwelling on, because both are the kind of thing that quietly\ninvalidates a result. First, an earlier draft of the cell above counted `izlenme` as a\nTurkish negator. It is not. `izlenmektedir` means the finding *is* seen and\n`izlenmemektedir` means it is not, so the loose pattern scored positives as negatives and\ninflated the negation rate. The number above uses a narrowed list, and the fact that this\nerror survived a first pass is the argument, not a footnote to it.\n\nSecond, look at the homoglyph count. Most Greek reports spell mu with `µ`, the MICRO SIGN\nat U+00B5, rather than `μ`, the Greek letter at U+03BC. The two are visually identical and\ncompare unequal. A pattern written by someone who typed Greek correctly matches almost\nnone of the Greek corpus, and it fails by returning zero matches, which reads as a negative\nlabel rather than an error.\n\nA regex list would have to encode negation scope, in every language present, with correct\nhandling of coordination, line breaks, and Unicode confusables. Building that is harder\nthan the imaging model. Extract labels with something that reads the sentence, and audit\nthe result against the gold 58, per label."},{"cell_type":"code","execution_count":null,"id":"bddc8f22","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Are any reports byte-identical across studies, and does that matter for folds?\"\"\"\n\nnorm = reports.str.strip().str.lower()\ncounts = norm.value_counts()\ndupes = counts[counts > 1]\n\nprint(f\"studies                      : {len(train):,}\")\nprint(f\"distinct report texts        : {norm.nunique():,}\")\nprint(f\"texts used by >1 study       : {len(dupes):,}\")\nprint(f\"studies sharing a text       : {int(dupes.sum()):,} \"\n      f\"({dupes.sum() / len(train):.1%})\")\nprint(f\"largest single group         : {int(dupes.max()) if len(dupes) else 0} studies\")\n\ngold_mask = train[LABELS].notna().all(axis=1)\ngold_dup = int(norm[gold_mask].isin(dupes.index).sum())\nprint(f\"gold studies with a shared text: {gold_dup} of {int(gold_mask.sum())}\")\n\ntop = dupes.head(6)\nshow_table(\n    pd.DataFrame({\n        \"studies\": top.to_numpy(),\n        \"language\": [lang[norm == t].mode().iat[0] for t in top.index],\n        \"chars\": [len(t) for t in top.index],\n        \"opening\": [t[:64].replace(\"\\n\", \" \") + \"...\" for t in top.index],\n    }).reset_index(drop=True),\n    caption=\"The largest groups are boilerplate normal reads, not copied studies.\",\n)"},{"cell_type":"markdown","id":"88b2b5ba","metadata":{},"source":"**Finding.** Report text is not unique per study. A small share of studies, a few percent,\nshare a byte-identical report with at least one other study, and the largest single group\nruns to dozens of studies all carrying the same normal read in the same language.\n\nThese are not duplicated studies. The images differ; the words a radiologist typed happen\nto be the same, which is what you would expect from a template used for an unremarkable\nknee. That makes them real training data rather than something to drop.\n\nIt does change how folds must be built. Any label extractor reads one text and hands the\nsame twelve-value target vector to every study in the group. Split that group across folds\nand the model is scored on targets it has already seen the exact source of, which flatters\nthe validation number without flattering the model. Keep a duplicate group whole inside one\nfold, the same way patients are kept whole, and the problem disappears."},{"cell_type":"markdown","id":"09ec81f3","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">4</span><span style=\"vertical-align:middle;\">Reports are training-time only</span></h2></div>\n\nThe finding most likely to save a reader a week."},{"cell_type":"code","execution_count":null,"id":"f1616779","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Show the train and test schemas, one above the other.\"\"\"\n\ntrain_cols = list(train.columns)\ntest_cols = list(test.columns)\nprint(\"train.csv columns:\", train_cols)\nprint(\"test.csv  columns:\", test_cols)\nprint(\"test.csv rows    :\", len(test), \"(placeholder rows; the real test set is hidden)\")\nprint()\nprint(\"columns in train.csv absent from test.csv:\",\n      [c for c in train_cols if c not in test_cols])\n\nfig, axes = plt.subplots(2, 1, figsize=(12.8, 8.0),\n                         gridspec_kw={\"hspace\": 0.34})\npanels = [\n    (\"train.csv\", axes[0],\n     [(\"StudyInstanceUID\", \"present\"),\n      (\"Report\", \"present\"),\n      (\"12 label columns\", \"present on 58 rows\")]),\n    (\"test.csv\", axes[1],\n     [(\"StudyInstanceUID\", \"present\"),\n      (\"Report\", \"ABSENT\"),\n      (\"12 label columns\", \"ABSENT\")]),\n]\nfor title, ax, rows_ in panels:\n    ax.axis(\"off\")\n    ax.set_xlim(0, 1)\n    ax.set_ylim(0, 1)\n    ax.text(0.02, 0.94, title, fontsize=16, fontweight=\"bold\", color=INK,\n            family=\"monospace\")\n    y = 0.74\n    for name, state in rows_:\n        absent = state == \"ABSENT\"\n        # PRIMARY_DEEP, not PRIMARY: the row label is drawn in SURFACE on top of\n        # this fill, and SURFACE on PRIMARY is 4.3:1, just under the AA floor.\n        ax.add_patch(plt.Rectangle((0.02, y - 0.09), 0.94, 0.17,\n                                   transform=ax.transAxes,\n                                   facecolor=SURFACE if absent else PRIMARY_DEEP,\n                                   edgecolor=NEUTRAL if absent else \"none\",\n                                   linewidth=1.6, linestyle=(0, (4, 3)) if absent else \"-\",\n                                   zorder=2))\n        ax.text(0.05, y - 0.005, name, transform=ax.transAxes, fontsize=15,\n                color=MUTED if absent else ink_on(PRIMARY_DEEP),\n                fontweight=\"bold\" if not absent else \"normal\", va=\"center\",\n                family=\"monospace\")\n        ax.text(0.93, y - 0.005, state, transform=ax.transAxes, fontsize=12,\n                color=ACCENT_TEXT if absent else ink_on(PRIMARY_DEEP),\n                va=\"center\", ha=\"right\")\n        y -= 0.26\n    ax.text(0.02, 0.02, f\"{len(train):,} rows\" if title == \"train.csv\"\n            else f\"{len(test):,} placeholder rows, real test hidden\",\n            transform=ax.transAxes, fontsize=12, color=MUTED)\n\nfinish(fig, \"test.csv has one column. There is no report at inference time\",\n       detail=\"Text is a label source and an auxiliary training signal. It is never an \"\n              \"input to a prediction.\",\n       caption=\"Dashed outline marks a column that exists in train.csv and does not exist \"\n               \"in test.csv. Every box is labelled, so nothing here is carried by colour alone.\")"},{"cell_type":"markdown","id":"4ee6e6d8","metadata":{},"source":"**Finding.** \"Multimodal\" in the competition description means the model may *learn* from\ntext. At inference it sees pixels and DICOM headers. Any pipeline that reads report text to\nproduce a prediction scores nothing, and it will look excellent in cross-validation right\nup until submission, because the report contains the answer.\n\nThree designs survive this constraint. Use text only to build labels, then train a pure\nimaging model. Use text as an auxiliary target during training, for instance by distilling\na text encoder's representation into the image encoder, and drop the text branch at\ninference. Or use text to build a curriculum, weighting studies by how confident the\nextraction was. All three are legitimate. Reading the report at predict time is not.\n\nThere is a related trap worth naming. If you build a report-derived validation set, the\nlabels on it came from the report, so a model that has seen that report during training has\nan advantage that will not transfer. Keep the split at the patient level and keep the\nextraction step out of the fold, exactly as you would for target encoding."},{"cell_type":"markdown","id":"e0f9ed05","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">5</span><span style=\"vertical-align:middle;\">Series structure</span></h2></div>\n\nFive or six series per study, three planes, and two provided protocol flags."},{"cell_type":"code","execution_count":null,"id":"6103bb7c","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Series per study, plane mix, and the two protocol flags.\"\"\"\n\nspc = train_series.groupby(\"StudyInstanceUID\").size()\nplanes = train_series[\"Anatomical_Plane\"].value_counts()\n\nprint(\"series per study:\", spc.describe()[[\"mean\", \"std\", \"min\", \"50%\", \"max\"]].round(2).to_dict())\nprint(\"studies covered :\", spc.size, \"of\", train[\"StudyInstanceUID\"].nunique())\nprint()\nprint(planes.to_string())\nprint()\nprint(\"every study has all three planes:\",\n      bool((train_series.groupby(\"StudyInstanceUID\")[\"Anatomical_Plane\"]\n            .nunique() == 3).all()))\nprint(\"plane counts per study (min across studies):\",\n      train_series.groupby(\"StudyInstanceUID\")[\"Anatomical_Plane\"].nunique().min())\n\nfig, axes = plt.subplots(2, 1, figsize=(13.5, 9.4),\n                         gridspec_kw={\"height_ratios\": [1.30, 1.0], \"hspace\": 0.34})\n\nax0 = axes[0]\nc = spc.value_counts().sort_index()\nbars = ax0.bar(c.index, c.values, width=0.72, color=PRIMARY, zorder=3)\nmode_k = int(c.idxmax())\nfor b, k in zip(bars, c.index):\n    if k == mode_k:\n        b.set_color(ACCENT)\nfor k, v in c.items():\n    if v > len(spc) * 0.01:\n        ax0.text(k, v + len(spc) * 0.012, f\"{v:,}\", ha=\"center\", va=\"bottom\",\n                 fontsize=11.5, color=INK_2)\nax0.set_ylim(0, c.max() * 1.12)\nax0.set_xticks(list(c.index))\nax0.set_xlabel(\"series in the study\")\nax0.set_ylabel(\"studies\")\nstyle(ax0, \"y\")\nax0.set_title(f\"Most studies have exactly {mode_k} series\", loc=\"left\",\n              fontsize=13.5, color=INK, pad=10)\n\nax1 = axes[1]\npv = planes.to_numpy()\nhbar(ax1, planes.index.tolist(), pv, [PRIMARY] * len(planes),\n     annot=[f\"{v:,}   {100 * v / len(train_series):.0f}%\" for v in pv], xpad=1.34)\nax1.set_xlabel(\"series\")\nax1.set_title(\"Sagittal is the most acquired plane\", loc=\"left\", fontsize=13.5,\n              color=INK, pad=10)\n\nfinish(fig, \"A study is five or six series, and all three planes are always present\",\n       headroom=0.38,\n       detail=f\"Median {int(spc.median())} series per study, range {int(spc.min())} to \"\n              f\"{int(spc.max())}. Sagittal {planes.get('Sagittal', 0):,}, \"\n              f\"Coronal {planes.get('Coronal', 0):,}, Axial {planes.get('Axial', 0):,}.\",\n       caption=\"Top: distribution over the 4,407 training studies. Bottom: plane mix \"\n               \"over all 24,371 series.\")"},{"cell_type":"markdown","id":"6a1737af","metadata":{},"source":"**Finding.** Series count varies from three to fourteen, so a model cannot assume a fixed\ninput shape at the study level. Something has to pool a variable-length set of series into\none prediction, and the obvious choices differ in how much they can learn: mean or max over\nper-series logits, attention over series embeddings, or a plane-conditioned head that\nconsumes exactly one sagittal, one coronal and one axial series. The last one is attractive\nbecause the plane composition is stable, and it is the reason the plane check above matters."},{"cell_type":"code","execution_count":null,"id":"4364add3","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Are Fluid_Sensitive and Fat_Suppression the same column? Answer, do not ask.\"\"\"\n\nfs = train_series[\"Fluid_Sensitive\"]\nfat = train_series[\"Fat_Suppression\"]\ndisagree = int((fs != fat).sum())\nidentical = disagree == 0\n\nprint(\"Fluid_Sensitive value counts:\", fs.value_counts().to_dict())\nprint(\"Fat_Suppression value counts:\", fat.value_counts().to_dict())\nprint()\nprint(pd.crosstab(fs, fat, rownames=[\"Fluid_Sensitive\"], colnames=[\"Fat_Suppression\"]))\nprint()\nprint(f\"rows where the two disagree : {disagree} of {len(train_series):,}\")\nprint(f\"IDENTICAL COLUMNS           : {identical}\")\n\nfs_te = test_series[\"Fluid_Sensitive\"]\nfat_te = test_series[\"Fat_Suppression\"]\nprint(f\"same check on test_series   : {int((fs_te != fat_te).sum())} disagreements \"\n      f\"of {len(test_series):,} rows\")\n\nfig, ax = plt.subplots(figsize=(14.0, 5.0))\nct = pd.crosstab(train_series[\"Anatomical_Plane\"], fs)\norder = [\"Sagittal\", \"Coronal\", \"Axial\"]\nct = ct.reindex([p for p in order if p in ct.index])\ny = np.arange(len(ct))[::-1]\nleft = np.zeros(len(ct))\n# PRIMARY_DEEP for the segment that carries a number inside it.\nseg_cols = [NEUTRAL, PRIMARY_DEEP]\nseg_names = [\"flag = 0\", \"flag = 1\"]\nfor j, col in enumerate(ct.columns):\n    w = ct[col].to_numpy(dtype=float)\n    ax.barh(y, w, left=left, height=0.66, color=seg_cols[j], zorder=3,\n            edgecolor=SURFACE, linewidth=2.0)\n    for yy, ww, ll in zip(y, w, left):\n        if ww > ct.to_numpy().sum() * 0.02:\n            ax.text(ll + ww / 2, yy, f\"{int(ww):,}\", ha=\"center\", va=\"center\",\n                    fontsize=12, fontweight=\"bold\", color=ink_on(seg_cols[j]))\n    left = left + w\nax.set_yticks(y)\nax.set_yticklabels(ct.index.tolist(), fontsize=12.5, color=INK)\nax.set_xlabel(\"series\")\nstyle(ax, \"x\")\nax.spines[\"left\"].set_visible(False)\nhandles = [Line2D([0], [0], marker=\"s\", linestyle=\"none\", markersize=11, label=n,\n                  markerfacecolor=c, markeredgecolor=\"none\")\n           for n, c in zip(seg_names, seg_cols)]\nax.legend(handles=handles, loc=\"lower right\", labelcolor=INK_2, borderaxespad=0.8)\n\nfinish(fig, \"Fluid_Sensitive and Fat_Suppression are the same column, byte for byte\",\n       detail=f\"{disagree} disagreements across {len(train_series):,} training series. \"\n              f\"Treat them as one binary protocol flag, not two features.\",\n       caption=\"Bars break the single flag down by plane. Axial series are overwhelmingly \"\n               \"flagged 1; sagittal series are close to an even split.\")"},{"cell_type":"markdown","id":"b06e20a5","metadata":{},"source":"**Finding.** They are identical. Zero disagreements in 24,371 training series and zero in\nthe test series file. Whatever the host intended, what shipped is one binary flag under two\nnames, so a feature set that uses both is duplicating a column and any importance ranking\nwill split its credit across the pair.\n\nThe flag is not distributed evenly across planes. Axial series are flagged 1 four times as\noften as flagged 0, while sagittal series split almost evenly. So the flag and the plane\ncarry overlapping information, and a plane-conditioned model already knows some of what the\nflag would tell it. The DICOM `SeriesDescription` and `ScanningSequence` tags give an\nindependent read on the same protocol question, and section 6 shows they are available at\ntest time."},{"cell_type":"markdown","id":"9e7678a9","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">6</span><span style=\"vertical-align:middle;\">DICOM metadata</span></h2></div>\n\nHeaders are read with `stop_before_pixels=True` from a bounded sample of series. Nothing\nhere touches pixel data, so it is cheap even on the full mount."},{"cell_type":"code","execution_count":null,"id":"92c2274b","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Discover series and read a bounded sample of headers.\"\"\"\n\nimport pydicom\nfrom pydicom.multival import MultiValue\n\n\ndef discover_series(max_series):\n    \"\"\"Find up to ``max_series`` DICOM series without walking every slice.\n\n    Prefers the nested ``train_series/<study>/<series>/*.dcm`` layout on the\n    Kaggle mount. Falls back to a bounded walk of whatever DICOM directories\n    exist locally, grouping files by the SeriesInstanceUID in their header.\n\n    Args:\n        max_series: Hard cap on the number of series returned.\n\n    Returns:\n        A list of ``(study_uid, series_uid, [file paths])`` tuples, and a string\n        naming the discovery mode used.\n    \"\"\"\n    out = []\n    for base in (TRAIN_DCM_DIR, TEST_DCM_DIR):\n        if not base.is_dir():\n            continue\n        studies = sorted(p.name for p in base.iterdir() if p.is_dir())\n        if not studies:\n            continue\n        k = min(N_HEADER_STUDIES, len(studies))\n        pick = np.random.default_rng(SEED).choice(len(studies), size=k, replace=False)\n        for i in sorted(pick):\n            sdir = base / studies[i]\n            for ser in sorted(p for p in sdir.iterdir() if p.is_dir()):\n                files = sorted(str(f) for f in ser.glob(\"*.dcm\"))\n                if files:\n                    out.append((studies[i], ser.name, files))\n                if len(out) >= max_series:\n                    return out, f\"nested layout under {base.name}\"\n        if out:\n            return out, f\"nested layout under {base.name}\"\n\n    # Fallback: a flat or unknown local layout. Bounded, then grouped by header.\n    files = []\n    for d in [TRAIN_DCM_DIR, TEST_DCM_DIR, *EXTRA_DCM_DIRS]:\n        if not d.is_dir():\n            continue\n        for dirpath, _, filenames in os.walk(d):\n            for fn in sorted(filenames):\n                if fn.lower().endswith(\".dcm\"):\n                    files.append(os.path.join(dirpath, fn))\n                    if len(files) >= MAX_FALLBACK_FILES:\n                        break\n            if len(files) >= MAX_FALLBACK_FILES:\n                break\n        if len(files) >= MAX_FALLBACK_FILES:\n            break\n    groups = collections.defaultdict(list)\n    for f in sorted(files):\n        try:\n            d = pydicom.dcmread(f, stop_before_pixels=True, specific_tags=[\n                \"StudyInstanceUID\", \"SeriesInstanceUID\"])\n            groups[(str(d.StudyInstanceUID), str(d.SeriesInstanceUID))].append(f)\n        except Exception:\n            continue\n    out = [(a, b, sorted(v)) for (a, b), v in sorted(groups.items())][:max_series]\n    return out, \"bounded walk, grouped by SeriesInstanceUID\"\n\n\nseries_index, discovery_mode = discover_series(MAX_HEADER_SERIES)\nn_files_seen = sum(len(f) for _, _, f in series_index)\nprint(f\"discovery mode : {discovery_mode}\")\nprint(f\"series found   : {len(series_index):,}\")\nprint(f\"slices indexed : {n_files_seen:,}\")\nif not series_index:\n    print(\"\\nNO DICOM FILES FOUND. Sections 6 and 7 will report that and continue.\")"},{"cell_type":"code","execution_count":null,"id":"13b06856","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Read one header per discovered series into a dataframe.\"\"\"\n\nHEADER_TAGS = [\n    \"PatientID\", \"PatientSex\", \"Laterality\", \"BodyPartExamined\",\n    \"Manufacturer\", \"ManufacturerModelName\", \"MagneticFieldStrength\",\n    \"SoftwareVersions\", \"ReceiveCoilName\", \"SeriesDescription\",\n    \"ScanningSequence\", \"SequenceVariant\", \"MRAcquisitionType\",\n    \"RepetitionTime\", \"EchoTime\", \"InversionTime\", \"FlipAngle\",\n    \"EchoTrainLength\", \"SliceThickness\", \"SpacingBetweenSlices\",\n    \"PixelSpacing\", \"Rows\", \"Columns\", \"BitsStored\", \"PhotometricInterpretation\",\n    \"WindowCenter\", \"WindowWidth\", \"RescaleSlope\", \"RescaleIntercept\",\n    \"InstanceNumber\", \"SeriesNumber\", \"ImageOrientationPatient\",\n]\n\n\ndef plane_from_iop(iop):\n    \"\"\"Derive the acquisition plane from ImageOrientationPatient.\n\n    The slice normal is the cross product of the two row/column direction\n    cosines; whichever patient axis it aligns with names the plane.\n\n    Args:\n        iop: Six-element ImageOrientationPatient value, or ``None``.\n\n    Returns:\n        ``\"Sagittal\"``, ``\"Coronal\"``, ``\"Axial\"`` or ``\"unknown\"``.\n    \"\"\"\n    try:\n        v = np.asarray([float(x) for x in iop], dtype=float)\n        n = np.abs(np.cross(v[:3], v[3:]))\n        return [\"Sagittal\", \"Coronal\", \"Axial\"][int(np.argmax(n))]\n    except Exception:\n        return \"unknown\"\n\n\nrows = []\nheader_failures = []\nfor study_uid, series_uid, files in series_index:\n    try:\n        ds = pydicom.dcmread(files[len(files) // 2], stop_before_pixels=True)\n    except Exception as exc:\n        header_failures.append((series_uid, type(exc).__name__, str(exc)[:90]))\n        continue\n    rec = {\"StudyInstanceUID\": study_uid, \"SeriesInstanceUID\": series_uid,\n           \"n_slices\": len(files), \"path\": files[len(files) // 2]}\n    for tag in HEADER_TAGS:\n        val = getattr(ds, tag, None)\n        if isinstance(val, MultiValue):\n            val = list(val)\n        rec[tag] = val\n    try:\n        rec[\"TransferSyntax\"] = str(ds.file_meta.TransferSyntaxUID.name)\n    except Exception:\n        rec[\"TransferSyntax\"] = \"unknown\"\n    rec[\"plane_from_iop\"] = plane_from_iop(rec.get(\"ImageOrientationPatient\"))\n    rows.append(rec)\n\nhdr = pd.DataFrame(rows)\nprint(f\"headers read    : {len(hdr):,}\")\nprint(f\"header failures : {len(header_failures)}\")\nfor f in header_failures[:5]:\n    print(\"   \", f)\n\nif len(hdr):\n    known = pd.concat([train_series, test_series], ignore_index=True)[\n        [\"SeriesInstanceUID\", \"Anatomical_Plane\", \"Fluid_Sensitive\"]]\n    hdr = hdr.merge(known, on=\"SeriesInstanceUID\", how=\"left\")\n    matched = hdr[\"Anatomical_Plane\"].notna() & (hdr[\"plane_from_iop\"] != \"unknown\")\n    if matched.any():\n        agree = (hdr.loc[matched, \"Anatomical_Plane\"]\n                 == hdr.loc[matched, \"plane_from_iop\"]).mean()\n        print(f\"\\nplane derived from ImageOrientationPatient agrees with the provided \"\n              f\"Anatomical_Plane on {100 * agree:.1f}% of {int(matched.sum()):,} \"\n              f\"matched series\")\n        print(\"(the derived plane comes from the direction cosines in the slice's own \"\n              \"header, so it is an independent check on the CSV rather than a copy of it)\")\n    display_cols = [\"Manufacturer\", \"ManufacturerModelName\", \"MagneticFieldStrength\",\n                    \"SeriesDescription\", \"Laterality\", \"PatientSex\", \"Rows\",\n                    \"SliceThickness\", \"TransferSyntax\", \"n_slices\"]\n    show_table(hdr[[c for c in display_cols if c in hdr.columns]].head(8))"},{"cell_type":"code","execution_count":null,"id":"9bcd4b70","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Scanner, protocol and encoding census over the sampled series.\"\"\"\n\nif not len(hdr):\n    print(\"No DICOM headers available at this data root; skipping the census charts.\")\nelse:\n    def top_counts(col, k=8):\n        \"\"\"Return the top ``k`` value counts of a header column as strings.\"\"\"\n        s = hdr[col].dropna().astype(str)\n        s = s[s.str.strip() != \"\"]\n        return s.value_counts().head(k)\n\n    # State what the sample actually contains rather than asserting variety the\n    # panels may not show: on a small local mirror every facet has one value.\n    def n_distinct(col):\n        \"\"\"Count distinct non-empty values of a header column.\"\"\"\n        if col not in hdr.columns:\n            return 0\n        s = hdr[col].dropna().astype(str)\n        return int(s[s.str.strip() != \"\"].nunique())\n\n    def census_pair(panels, finding, caption_tail):\n        \"\"\"Draw two stacked category panels and close the figure out.\n\n        Args:\n            panels: Two ``(column, title)`` pairs, drawn top then bottom.\n            finding: The headline for this figure.\n            caption_tail: Sentence appended to the shared provenance caption.\n        \"\"\"\n        fig, axes = plt.subplots(2, 1, figsize=(13.5, 9.8),\n                                 gridspec_kw={\"hspace\": 0.40})\n        for ax, (col, title) in zip(axes, panels):\n            if col not in hdr.columns:\n                ax.axis(\"off\")\n                continue\n            vc = top_counts(col, 8)\n            if not len(vc):\n                ax.axis(\"off\")\n                ax.text(0.5, 0.5, f\"{col}: no values\", ha=\"center\", va=\"center\",\n                        color=MUTED, transform=ax.transAxes)\n                continue\n            names = [n if len(n) <= 34 else n[:32] + \"…\" for n in vc.index.tolist()]\n            cols_ = [ACCENT if i == 0 else PRIMARY for i in range(len(vc))]\n            hbar(ax, names, vc.to_numpy(), cols_, fmt=\"{:,.0f}\")\n            # Reserve four bar slots even when the sample holds fewer values,\n            # padding downward so the bars stay top-aligned. On the full mount\n            # every facet fills all eight and this does nothing; on a one-vendor\n            # local mirror it stops a single bar swelling to half the panel.\n            ax.set_ylim(len(vc) - max(len(vc), 4) - 0.75, len(vc) - 0.25)\n            ax.set_title(title, loc=\"left\", fontsize=13.5, color=INK, pad=10)\n            ax.set_xlabel(\"series sampled\")\n        finish(fig, finding, headroom=0.38,\n               detail=f\"Census over {len(hdr):,} series sampled from \"\n                      f\"{hdr['StudyInstanceUID'].nunique():,} studies, one header \"\n                      f\"each. A small local mirror will show one value per facet; \"\n                      f\"the full 265 GB mount does not.\",\n               caption=f\"Bounded by N_HEADER_STUDIES={N_HEADER_STUDIES} and \"\n                       f\"MAX_HEADER_SERIES={MAX_HEADER_SERIES}, seed {SEED}. \"\n                       f\"Discovery mode: {discovery_mode}. Top 8 values per panel. \"\n                       + caption_tail)\n\n    n_vendor = n_distinct(\"Manufacturer\")\n    n_field = n_distinct(\"MagneticFieldStrength\")\n    n_syntax = n_distinct(\"TransferSyntax\")\n    n_model = n_distinct(\"ManufacturerModelName\")\n\n    census_pair(\n        [(\"Manufacturer\", \"Scanner vendor\"),\n         (\"ManufacturerModelName\", \"Scanner model\")],\n        f\"{n_vendor} vendor{'' if n_vendor == 1 else 's'} and {n_model} scanner \"\n        f\"model{'' if n_model == 1 else 's'} in this sample\",\n        \"Vendor and model are the closest thing the corpus has to a site label.\")\n\n    census_pair(\n        [(\"MagneticFieldStrength\", \"Field strength (T)\"),\n         (\"TransferSyntax\", \"Pixel encoding\")],\n        f\"{n_field} field strength{'' if n_field == 1 else 's'} and {n_syntax} \"\n        f\"pixel encoding{'' if n_syntax == 1 else 's'} in this sample\",\n        \"Pixel encoding decides which decoders a submission kernel has to ship.\")"},{"cell_type":"markdown","id":"89992a4b","metadata":{},"source":"**Finding.** The host describes \"a diverse international mix of imaging sites\". Vendor,\nmodel, field strength, coil and software version are all recorded in the header, which\nmakes them a usable proxy for the site. That cuts two ways. They are the right variables\nfor a domain-shift check, and they are exactly the shortcut a network will find if one site\nhappens to contribute most of the positives for a finding. Check whether any label's\nprevalence correlates with vendor before trusting a validation score.\n\nPixel encoding matters for a more practical reason. Transfer syntaxes are mixed across the\nfull corpus, so the decode path has to handle uncompressed Explicit VR, Implicit VR, JPEG\nLossless and JPEG 2000. A submission kernel that only decodes uncompressed data does not\nscore badly on the hidden test set, it fails. Verify the decoders in the submission\nenvironment rather than trusting whatever the interactive session happened to sample.\nKaggle's image includes the JPEG plugins; a self-built container may not."},{"cell_type":"code","execution_count":null,"id":"df3c78cd","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Patient grouping, laterality and sex. The fold-construction question.\"\"\"\n\nif not len(hdr):\n    print(\"No DICOM headers available; skipping the patient-grouping check.\")\nelse:\n    pid = hdr.dropna(subset=[\"PatientID\"])\n    per_patient = pid.groupby(\"PatientID\")[\"StudyInstanceUID\"].nunique()\n    multi = int((per_patient > 1).sum())\n    print(f\"distinct PatientIDs in the sample : {per_patient.size:,}\")\n    print(f\"studies in the sample             : {pid['StudyInstanceUID'].nunique():,}\")\n    print(f\"patients contributing >1 study    : {multi}\")\n    print(f\"studies per patient (max)         : {int(per_patient.max())}\")\n    print()\n    print(\"PatientID example (pseudonymised):\", pid[\"PatientID\"].iloc[0])\n    print()\n    print(\"Laterality:\", hdr[\"Laterality\"].dropna().astype(str).value_counts().to_dict())\n    print(\"PatientSex:\", hdr[\"PatientSex\"].dropna().astype(str).value_counts().to_dict())\n    print()\n    per_study_pid = pid.groupby(\"StudyInstanceUID\")[\"PatientID\"].nunique()\n    print(\"studies whose series disagree on PatientID:\",\n          int((per_study_pid > 1).sum()), \"of\", per_study_pid.size)\n\n    fig, axes = plt.subplots(2, 1, figsize=(13.0, 6.6),\n                             gridspec_kw={\"hspace\": 0.50})\n    lat = hdr[\"Laterality\"].dropna().astype(str).value_counts()\n    sex = hdr[\"PatientSex\"].dropna().astype(str).value_counts()\n    for ax, vc, title, xl in [(axes[0], lat, \"Which knee\", \"series\"),\n                              (axes[1], sex, \"Patient sex\", \"series\")]:\n        if len(vc):\n            hbar(ax, vc.index.tolist(), vc.to_numpy(), [PRIMARY] * len(vc))\n            ax.set_xlabel(xl)\n        else:\n            ax.axis(\"off\")\n            ax.text(0.5, 0.5, \"tag absent\", ha=\"center\", va=\"center\", color=MUTED,\n                    transform=ax.transAxes)\n        ax.set_title(title, loc=\"left\", fontsize=13.5, color=INK, pad=10)\n\n    finish(fig, \"Laterality and patient sex are in the DICOM header, and so is a patient key\",\n           headroom=0.38, footroom=0.16,\n           detail=\"PatientID is pseudonymised but stable, which makes it the grouping key \"\n                  \"for cross-validation.\",\n           caption=f\"Counted over the {len(hdr):,} sampled series. \"\n                   \"Laterality matters: a left-knee model must not learn the mirror image.\")"},{"cell_type":"markdown","id":"0a8f3acc","metadata":{},"source":"**Finding.** `PatientID` is present and pseudonymised; the cell above prints one, along\nwith the number of patients contributing more than one study in the sampled slice of the\ncorpus. Group folds by `PatientID`, not by `StudyInstanceUID`. A patient with two studies\nputs the same knee on both sides of a study-level split, and knee MRI is about as\npatient-specific as imaging gets: the same degenerative changes appear in both. Grouping\ncosts nothing when every patient turns out to have exactly one study, so build the\nstudy-to-patient map over the whole corpus and group on it either way.\n\n`Laterality` deserves attention for a different reason. Left and right knees are mirror\nimages, so either normalise by flipping all rights to lefts, or feed laterality to the model\nexplicitly. Doing neither forces the network to spend capacity learning a symmetry it was\nhanded for free. Both `Laterality` and `PatientSex` come from the DICOM header rather than\n`train.csv`, so both are available at test time."},{"cell_type":"code","execution_count":null,"id":"01ee5d03","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"What the SeriesDescription strings say about the protocol.\"\"\"\n\nif not len(hdr) or \"SeriesDescription\" not in hdr.columns:\n    print(\"SeriesDescription unavailable; skipping.\")\nelse:\n    desc = hdr[\"SeriesDescription\"].dropna().astype(str).str.lower()\n    print(f\"distinct SeriesDescription values in the sample: {desc.nunique():,} \"\n          f\"over {len(desc):,} series\")\n    print(\"\\nmost common raw values:\")\n    print(desc.value_counts().head(12).to_string())\n\n    tokens = collections.Counter()\n    for d in desc:\n        for t in re.split(r\"[^a-z0-9]+\", d):\n            if t and not t.isdigit() and len(t) > 1:\n                tokens[t] += 1\n    top = tokens.most_common(18)\n\n    fig, ax = plt.subplots(figsize=(14.0, 7.4))\n    PROTOCOL_HINTS = {\"fs\", \"stir\", \"spair\", \"sag\", \"cor\", \"tra\", \"ax\", \"t1\", \"t2\",\n                      \"pd\", \"dp\", \"tirm\", \"fatsat\", \"sagittal\", \"coronal\", \"axial\"}\n    names = [t for t, _ in top]\n    vals = [c for _, c in top]\n    cols_ = [ACCENT if n in PROTOCOL_HINTS else PRIMARY for n in names]\n    hbar(ax, names, vals, cols_, fmt=\"{:,.0f}\")\n    ax.set_xlabel(\"series containing the token\")\n    handles = [Line2D([0], [0], marker=\"s\", linestyle=\"none\", markersize=11, label=t,\n                      markerfacecolor=c, markeredgecolor=\"none\")\n               for t, c in [(\"plane or contrast token\", ACCENT), (\"other token\", PRIMARY)]]\n    ax.legend(handles=handles, loc=\"lower right\", labelcolor=INK_2, borderaxespad=0.8)\n\n    finish(fig, \"SeriesDescription encodes plane and contrast weighting in free text\",\n           detail=\"Tokens like sag, cor, tra, pd, t2 and fs recover the acquisition \"\n                  \"protocol independently of the provided CSV flags.\",\n           caption=f\"Tokenised on non-alphanumeric boundaries over {len(desc):,} sampled \"\n                   \"series; pure digits dropped. Vendor naming is inconsistent, so treat \"\n                   \"this as a signal to verify, not a parser to trust.\")"},{"cell_type":"markdown","id":"eb011047","metadata":{},"source":"**Finding.** `SeriesDescription` is vendor free text and not a controlled vocabulary, so\nparsing it is a heuristic. It is still useful as a cross-check on `Anatomical_Plane` and on\nthe single protocol flag from section 5, and it is available at test time. The stronger\nindependent check is geometric: the plane derived from `ImageOrientationPatient` is\ncomputed from direction cosines rather than from a string, and the agreement rate against\nthe provided `Anatomical_Plane` is printed above.\n\n**Which tags are usable at inference.** All of them. The test DICOMs come from the same\nde-identification pipeline as the training DICOMs, so `PatientID`, `Laterality`,\n`PatientSex`, `Manufacturer`, `MagneticFieldStrength`, `SeriesDescription`, the geometry\ntags and the windowing tags are present at test time exactly as they are here. The only\nthing missing at test time is the report. That asymmetry is worth stating in reverse: DICOM\nmetadata is fair game as model input, free text is not.\n\nOne caveat on `PatientID` at test time. It is legitimate as a *feature* only if you are\ncomfortable with what it implies, and it is far more valuable as a grouping key for\ntest-time aggregation: if two test studies share a patient, their predictions are not\nindependent."},{"cell_type":"markdown","id":"01bcf0f8","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">7</span><span style=\"vertical-align:middle;\">Pixels</span></h2></div>\n\nA montage across the three planes with correct windowing, then the reason per-series\nnormalisation is not optional."},{"cell_type":"code","execution_count":null,"id":"7b7996cd","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Windowing helpers, then a montage across the three acquisition planes.\"\"\"\n\n\ndef load_windowed(path):\n    \"\"\"Read one slice and map it to display range with the DICOM's own window.\n\n    Applies RescaleSlope and RescaleIntercept, then WindowCenter and WindowWidth\n    when the header carries them, falling back to the 1st-99th percentile.\n    MONOCHROME1 is inverted so bright always means high signal.\n\n    Args:\n        path: Path to a DICOM file.\n\n    Returns:\n        A ``(H, W)`` float array in ``[0, 1]``, and the dataset, or\n        ``(None, None)`` if the pixel data could not be decoded.\n    \"\"\"\n    try:\n        ds = pydicom.dcmread(path)\n        arr = ds.pixel_array.astype(np.float32)\n    except Exception:\n        return None, None\n    slope = float(getattr(ds, \"RescaleSlope\", 1.0) or 1.0)\n    inter = float(getattr(ds, \"RescaleIntercept\", 0.0) or 0.0)\n    arr = arr * slope + inter\n    wc, ww = getattr(ds, \"WindowCenter\", None), getattr(ds, \"WindowWidth\", None)\n    try:\n        wc = float(wc[0] if isinstance(wc, MultiValue) else wc)\n        ww = float(ww[0] if isinstance(ww, MultiValue) else ww)\n    except (TypeError, ValueError):\n        wc = ww = None\n    if wc is None or ww is None or ww <= 0:\n        lo_, hi_ = np.percentile(arr, [1, 99])\n    else:\n        lo_, hi_ = wc - ww / 2.0, wc + ww / 2.0\n    if hi_ <= lo_:\n        lo_, hi_ = float(arr.min()), float(arr.max()) or 1.0\n    out = np.clip((arr - lo_) / (hi_ - lo_), 0.0, 1.0)\n    if str(getattr(ds, \"PhotometricInterpretation\", \"MONOCHROME2\")) == \"MONOCHROME1\":\n        out = 1.0 - out\n    return out, ds\n\n\nPLANE_ORDER = [\"Sagittal\", \"Coronal\", \"Axial\"]\npixel_failures = []\n\nif not len(hdr):\n    print(\"No DICOM data at this root; skipping the montage.\")\n    montage = {}\nelse:\n    files_by_series = {s: f for _, s, f in series_index}\n    # Prefer the plane computed from the slice's own direction cosines: the\n    # panel label then describes the pixels actually on screen, whatever the\n    # CSV says.\n    plane_col = hdr[\"plane_from_iop\"].where(hdr[\"plane_from_iop\"] != \"unknown\",\n                                            hdr[\"Anatomical_Plane\"])\n    montage = {}\n    for plane in PLANE_ORDER:\n        cand = hdr.index[plane_col == plane].tolist()\n        if not cand:\n            continue\n        cand = cand[:N_MONTAGE_CANDIDATES]\n        for i in cand:\n            row = hdr.loc[i]\n            files = files_by_series.get(row[\"SeriesInstanceUID\"], [])\n            if len(files) < 2:\n                continue\n            k = min(N_MONTAGE_SLICES, len(files))\n            idx = (np.linspace(len(files) * 0.25, len(files) * 0.75, k)\n                   .round().astype(int).clip(0, len(files) - 1))\n            # Short series round two positions onto the same slice, which would\n            # put the identical image in two panels. Keep first occurrences.\n            seen, keep = set(), []\n            for j in idx.tolist():\n                if j not in seen:\n                    seen.add(j)\n                    keep.append(j)\n            picks = [files[j] for j in keep]\n            imgs = []\n            for p in picks:\n                img, _ = load_windowed(p)\n                if img is None:\n                    pixel_failures.append(p)\n                else:\n                    imgs.append(img)\n            if imgs:\n                montage[plane] = (row, imgs)\n                break\n    print(\"planes with decodable pixel data:\", list(montage))\n    print(\"slice decode failures:\", len(pixel_failures))"},{"cell_type":"code","execution_count":null,"id":"ba5aef3d","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Render the montage: three planes down, slices across, one greyscale.\"\"\"\n\nif not montage:\n    print(\"No decodable pixel data available at this data root. \"\n          \"On the Kaggle mount this cell renders a 3 x \"\n          f\"{N_MONTAGE_SLICES} montage.\")\nelse:\n    ncol = max(len(v[1]) for v in montage.values())\n    nrow = len(montage)\n    # The montage is the visual centrepiece of the notebook, so it is the one\n    # grid that gets the full column width. Width is capped at 14.4 inches so a\n    # five-slice run cannot push it past what Kaggle's column can show.\n    cell_w = min(3.45, 14.4 / max(ncol, 1))\n    fig, axes = plt.subplots(nrow, ncol, figsize=(cell_w * ncol, 3.65 * nrow),\n                             squeeze=False)\n    stroke = [pe.withStroke(linewidth=2.8, foreground=\"#000000\")]\n    for r, plane in enumerate([p for p in PLANE_ORDER if p in montage]):\n        row, imgs = montage[plane]\n        desc = str(row.get(\"SeriesDescription\") or \"series description absent\")\n        for c in range(ncol):\n            ax = axes[r][c]\n            ax.set_xticks([])\n            ax.set_yticks([])\n            for s in ax.spines.values():\n                s.set_visible(False)\n            if c >= len(imgs):\n                ax.axis(\"off\")\n                continue\n            ax.imshow(imgs[c], cmap=\"gray\", vmin=0.0, vmax=1.0,\n                      interpolation=\"bilinear\")\n            if c == 0:\n                ax.text(0.035, 0.960, plane, transform=ax.transAxes, color=\"#ffffff\",\n                        fontsize=15, fontweight=\"bold\", va=\"top\", ha=\"left\",\n                        path_effects=stroke)\n                ax.text(0.035, 0.885, desc[:26], transform=ax.transAxes,\n                        color=\"#ffffff\", fontsize=11.5, va=\"top\", ha=\"left\",\n                        path_effects=stroke)\n            ax.text(0.965, 0.035, f\"slice {c + 1}\", transform=ax.transAxes,\n                    color=\"#ffffff\", fontsize=11, va=\"bottom\", ha=\"right\",\n                    path_effects=stroke)\n    top_in = 1.18\n    bot_in = 0.50\n    h_in = fig.get_size_inches()[1]\n    fig.subplots_adjust(left=0.004, right=0.996, wspace=0.010, hspace=0.014,\n                        top=1.0 - top_in / h_in, bottom=bot_in / h_in)\n\n    det = (f\"One series per plane, {ncol} slices each, drawn from the middle half of \"\n           f\"the stack. Windowed with the header's own WindowCenter/WindowWidth.\")\n    finish(fig, \"Three planes, three different views of the same joint\", detail=det,\n           tight=False,\n           caption=f\"RescaleSlope/Intercept applied before windowing; MONOCHROME1 \"\n                   f\"inverted. {len(pixel_failures)} slice decode failures in this sample. \"\n                   f\"Sequences shown: \"\n                   + \"; \".join(f\"{p}: {str(montage[p][0].get('SeriesDescription'))}\"\n                               for p in montage))"},{"cell_type":"markdown","id":"a151cb6a","metadata":{},"source":"**Finding.** The three planes are not redundant views of one thing, they are the reason a\nradiologist reads a knee in three planes. Cruciate ligaments run obliquely and are read\nsagittally. Collateral ligaments and the meniscal body are read coronally. Patellar\ncartilage and the retinacula are read axially. Since the twelve findings map unevenly onto\nthe planes, a model that pools all series identically is throwing away the structure the\nprotocol was designed around. Condition on the plane.\n\nWindowing is not cosmetic. Applying `RescaleSlope` and `RescaleIntercept` before the window\nis required, and skipping the window entirely leaves most of the dynamic range unused on\nthese images."},{"cell_type":"code","execution_count":null,"id":"9a6ae557","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Do intensities compare across series? Small multiples, same axis.\"\"\"\n\nif not len(hdr):\n    print(\"No DICOM data; skipping the intensity comparison.\")\nelse:\n    files_by_series = {s: f for _, s, f in series_index}\n    picked, stats = [], []\n    for i in hdr.index:\n        if len(picked) >= N_INTENSITY_SERIES:\n            break\n        row = hdr.loc[i]\n        files = files_by_series.get(row[\"SeriesInstanceUID\"], [])\n        if not files:\n            continue\n        try:\n            ds = pydicom.dcmread(files[len(files) // 2])\n            arr = ds.pixel_array.astype(np.float32)\n            arr = arr * float(getattr(ds, \"RescaleSlope\", 1.0) or 1.0) \\\n                + float(getattr(ds, \"RescaleIntercept\", 0.0) or 0.0)\n        except Exception as exc:\n            pixel_failures.append((files[0], type(exc).__name__))\n            continue\n        picked.append((row, arr))\n        p1, p50, p99 = np.percentile(arr, [1, 50, 99])\n        stats.append({\n            \"SeriesDescription\": str(row.get(\"SeriesDescription\"))[:24],\n            \"plane\": str(row.get(\"plane_from_iop\") or row.get(\"Anatomical_Plane\")),\n            \"p1\": round(float(p1), 1), \"median\": round(float(p50), 1),\n            \"p99\": round(float(p99), 1), \"max\": round(float(arr.max()), 1),\n        })\n\n    if not picked:\n        print(\"No decodable series for the intensity comparison.\")\n    else:\n        hi_all = max(np.percentile(a, 99.5) for _, a in picked)\n        n = len(picked)\n        ncol = min(3, n)\n        nrow = int(np.ceil(n / ncol))\n        fig, axes = plt.subplots(nrow, ncol, figsize=(4.8 * ncol, 3.7 * nrow),\n                                 squeeze=False, sharex=True)\n        for k, (row, arr) in enumerate(picked):\n            ax = axes[k // ncol][k % ncol]\n            ax.hist(arr.ravel(), bins=90, range=(0, hi_all), color=PRIMARY,\n                    edgecolor=\"none\", zorder=3, log=True)\n            # Headroom above the tallest bin, so the median label sits in clear\n            # space. Without it a series whose median lands on its own peak\n            # writes the label straight over the bars.\n            ax.set_ylim(top=ax.get_ylim()[1] * 8.0)\n            med = float(np.median(arr))\n            ax.axvline(med, color=ACCENT, linewidth=1.8, zorder=4)\n            ax.text(med + hi_all * 0.02, ax.get_ylim()[1] * 0.30,\n                    f\"median {med:,.0f}\", color=ACCENT_TEXT, fontsize=11,\n                    va=\"center\")\n            ttl = str(row.get(\"SeriesDescription\") or \"series\")[:24]\n            ax.set_title(f\"{ttl}  ·  {row.get('plane_from_iop') or row.get('Anatomical_Plane')}\",\n                         loc=\"left\", fontsize=12.5, color=INK)\n            style(ax, \"y\")\n            ax.set_yticklabels([])\n        for k in range(n, nrow * ncol):\n            axes[k // ncol][k % ncol].axis(\"off\")\n        for c in range(ncol):\n            axes[nrow - 1][c].set_xlabel(\"pixel value after rescale\")\n\n        meds = [s[\"median\"] for s in stats]\n        n_studies_shown = len({r[\"StudyInstanceUID\"] for r, _ in picked})\n        finish(fig, \"Raw intensity is not comparable across series\",\n               headroom=0.40,\n               detail=f\"Medians run from {min(meds):,.0f} to {max(meds):,.0f} across \"\n                      f\"these {n} series from {n_studies_shown} \"\n                      f\"stud{'y' if n_studies_shown == 1 else 'ies'}. MR signal has no \"\n                      f\"absolute unit.\",\n               caption=f\"Log-scaled counts, shared x-axis to {hi_all:,.0f}. One mid-stack \"\n                       f\"slice per series. Compare with CT, where Hounsfield units are \"\n                       f\"absolute and a fixed window is meaningful everywhere.\")\n\n        show_table(pd.DataFrame(stats),\n                   caption=\"Per-series percentiles after RescaleSlope/Intercept.\")"},{"cell_type":"markdown","id":"fa30267b","metadata":{},"source":"**Finding.** MR intensity is arbitrary. There is no Hounsfield-unit equivalent, so the same\ntissue takes a different number on a different sequence, a different coil, or a different\nday. A fixed global normalisation applied across this corpus will hand the network a\nbrightness offset that tracks the acquisition site, which is a shortcut correlated with\nwhichever findings that site over-represents.\n\nNormalise per series. The simple and hard-to-beat choice scales to the 1st and 99th\npercentile of the series, computed over the whole volume rather than per slice, so\nslices keep their relative contrast. Z-scoring on the non-background voxels is the other\nreasonable option. Whichever you pick, compute it inside the fold and apply the identical\ntransform at test time."},{"cell_type":"markdown","id":"b87bb20e","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"display:inline-block;min-width:26px;height:26px;line-height:26px;text-align:center;background:#256abf;color:#ffffff;border-radius:6px;font-size:14px;font-weight:700;margin-right:12px;vertical-align:middle;\">8</span><span style=\"vertical-align:middle;\">What this means for modelling</span></h2></div>\n\nNine decisions that follow from what is above, and one thing worth measuring before\ncommitting to any of them.\n\n**Build the label extractor first, and treat it as the main model.** 4,349 reports have to\nbecome 52,188 binary decisions. Everything downstream is capped by how good those decisions\nare. A keyword or regex approach has to survive seven or more languages and a negation rate\naround three quarters on at least one finding; that is a harder engineering problem than\nthe imaging model, and it fails silently when it fails. Use something that reads the\nsentence.\n\n**Audit the extractor on the gold 58, per label, and report agreement with an interval.**\nThat is the one job 58 studies can do. Twelve agreement rates with honest intervals tell you\nwhich findings the extraction is reliable on and which need a different prompt or a\ndifferent rule. Do not compute a macro AUC on those 58 and use it to pick anything.\n\n**Rank models on a report-derived validation set, not the gold set.** Hold out several\nhundred studies, label them with the audited extractor, and rank there. Group folds by\n`PatientID`. A patient with two studies would otherwise land on both sides of the split,\nand knee anatomy is patient-specific. Build the study-to-patient map over the full corpus\nbefore splitting.\n\n**Never read report text at inference.** Text is a label source, an auxiliary training\ntarget, or a curriculum weight. It is not an input. A validation score that uses it will\nlook excellent and transfer to nothing.\n\n**Condition on the plane.** All three planes are present in every study, the composition is\nstable, and different findings live in different planes. A plane-conditioned head, or at\nminimum a plane embedding, uses structure the protocol already provides.\n\n**Handle a variable number of series.** Three to fourteen, median five. Aggregation across\nseries is a design decision that has to be made explicitly, and attention over series\nembeddings is the version that can learn which series matters for which finding.\n\n**Normalise per series, inside the fold.** Percentile scaling over the whole volume.\nGlobal normalisation encodes the acquisition site.\n\n**Use `Fluid_Sensitive` once.** It and `Fat_Suppression` are the same column, with zero\ndisagreements in 24,371 rows. Including both duplicates a feature and splits its importance.\n\n**Spend on the rare findings.** Macro averaging means MCL, with nine positives in the gold\nset, is worth exactly what Effusion is worth. The temptation is to optimise the average by\nimproving what is already good. Under a macro metric that is close to worthless.\n\n**And measure this first:** whether label prevalence correlates with scanner vendor or\nfield strength. The header census shows the site proxies are all present and varied. If one\nsite supplies most of the positives for a finding, a network will learn the site, and both\nyour validation and the public leaderboard will reward it right up until the private split."},{"cell_type":"markdown","id":"49e0a600","metadata":{},"source":"<div style=\"background:#fcfcfb;border:1px solid #e1e0d9;border-left:6px solid #eb6834;border-radius:7px;padding:14px 18px;margin:26px 0 14px 0;\"><h2 style=\"color:#0b0b0b;font-size:21px;margin:0;font-weight:700;letter-spacing:-.005em;\"><span style=\"vertical-align:middle;\">How this notebook was made</span></h2></div>\n\nThis notebook was generated by an agentic system rather than typed by hand. The system is\nClaude Code running a multi-agent workflow: parallel agents wrote the analysis sections\nwhile separate adversarial verifier agents executed every cell against a local mirror of\nthe Kaggle input mount, checked every printed number against a fact sheet built from the\nKaggle API, opened every figure and looked at it, then sent each defect back to the builder\nscript instead of patching the notebook. The verifier agents are the part that matters,\nbecause a notebook that has never been run is a draft.\n\nThe source of truth is a Python builder script, `notebooks/build_eda_notebook.py`, which\nconstructs the `.ipynb` with `nbformat`. The notebook is an artifact of that script. Review\nhappens on plain Python rather than on notebook JSON, and the notebook is regenerated\nrather than edited. Publication uses the Kaggle CLI, `kaggle kernels push`, and the\ncompanion baseline notebook is submitted to the competition the same way.\n\nThe competition facts this notebook is built on were verified from the Kaggle API and from\nthe data itself on 2026-08-05, the day the competition opened. Study and series counts,\nthe 58 labelled studies, the twelve label names and their submission order, the\nidentity of `Fluid_Sensitive` and `Fat_Suppression`, and the single-column `test.csv` are\nall recomputed at run time in the cells above rather than quoted from that check. The\nsystem has produced no score on this competition and this notebook claims none."},{"cell_type":"code","execution_count":null,"id":"5a3de6fe","metadata":{"_kg_hide-input":true,"jupyter":{"source_hidden":true}},"outputs":[],"source":"\"\"\"Environment and reproducibility footer.\"\"\"\n\nimport platform\n\nprint(\"python      \", platform.python_version())\nfor mod in [\"numpy\", \"pandas\", \"matplotlib\", \"scipy\", \"pydicom\"]:\n    try:\n        print(f\"{mod:12s}\", __import__(mod).__version__)\n    except Exception as exc:\n        print(f\"{mod:12s} unavailable ({type(exc).__name__})\")\nprint()\nprint(\"data root         \", ROOT)\nprint(\"seed              \", SEED)\nprint(\"header budget     \", f\"{N_HEADER_STUDIES} studies / {MAX_HEADER_SERIES} series\")\nprint(\"series indexed    \", len(series_index))\nprint(\"headers read      \", len(hdr))\nprint(\"header failures   \", len(header_failures))\nprint(\"pixel decode fails\", len(pixel_failures))"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.0"}},"nbformat":4,"nbformat_minor":5}