{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.10"}},"nbformat_minor":4,"nbformat":4,"cells":[{"id":"255a4680-e391-43fc-9682-cc2be113e282","cell_type":"markdown","source":"# RSNA Knee Abnormality Detection — EDA & How Much Can We Trust the Reports?\n\nQuick context before diving in: this is a code competition, the test set has **no report text** (only `train.csv` has `Report`, `test.csv` doesn't). So the reports are only useful as a way to *manufacture more training labels* — not as a model input at inference time.\n\nOnly a small subset of studies come with the 12 ground-truth labels. Everything else only has a free-text radiology report. So the real question this notebook tries to answer with numbers, not vibes:\n\n> **If I write a keyword-based labeler and run it on the reports, how close does it get to the real labels?**\n\nIf it's decent, we can pseudo-label the huge unlabeled chunk and get way more training signal for free. If it's garbage, we need a smarter (LLM-based) extractor before trusting it.\n\nSpoiler from running this once already: it's not decent yet for every label. Keep reading, the per-label table in section 4 tells you exactly where it breaks and why.\n","metadata":{}},{"id":"2d9b3574-0039-454a-96a5-ade2ac5d2bc1","cell_type":"code","source":"import os\nimport glob\nimport json\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nsns.set_theme(style=\"whitegrid\")\nplt.rcParams[\"figure.dpi\"] = 110\n\nDATA_DIR = \"/kaggle/input/competitions/rsna-knee-abnormality-detection\"\nprint(os.listdir(DATA_DIR))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:28.23272Z","iopub.execute_input":"2026-08-05T18:25:28.23342Z","iopub.status.idle":"2026-08-05T18:25:28.268708Z","shell.execute_reply.started":"2026-08-05T18:25:28.233384Z","shell.execute_reply":"2026-08-05T18:25:28.267836Z"}},"outputs":[],"execution_count":null},{"id":"51fbea43-da04-4e66-bbbf-776fc853b85b","cell_type":"code","source":"# Compressed DICOM (JPEG2000 / JPEG Lossless) will fail to decode with plain pydicom.\n# The dataset card explicitly says transfer syntaxes are mixed, so install the decoders now\n# or you'll get cryptic \"unable to decode pixel data\" errors halfway through a loop.\n!pip install -q pylibjpeg pylibjpeg-libjpeg pylibjpeg-openjpeg python-gdcm\nimport pydicom\nprint(pydicom.__version__)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:28.38195Z","iopub.execute_input":"2026-08-05T18:25:28.382685Z","iopub.status.idle":"2026-08-05T18:25:32.379425Z","shell.execute_reply.started":"2026-08-05T18:25:28.382649Z","shell.execute_reply":"2026-08-05T18:25:32.378426Z"}},"outputs":[],"execution_count":null},{"id":"3a19f0b4-e189-4154-9c91-104a7ee4538b","cell_type":"markdown","source":"## 1. Loading the tables","metadata":{}},{"id":"900ea608-ccb4-458c-991b-4770dd8bad66","cell_type":"code","source":"train = pd.read_csv(f\"{DATA_DIR}/train.csv\")\ntrain_series = pd.read_csv(f\"{DATA_DIR}/train_series.csv\")\n\nLABEL_COLS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n              \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n              \"Contusion\", \"Fracture\"]\n\nprint(train.shape, train_series.shape)\ntrain.head(3)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:32.381522Z","iopub.execute_input":"2026-08-05T18:25:32.381823Z","iopub.status.idle":"2026-08-05T18:25:32.57017Z","shell.execute_reply.started":"2026-08-05T18:25:32.381788Z","shell.execute_reply":"2026-08-05T18:25:32.569485Z"}},"outputs":[],"execution_count":null},{"id":"c1404d66-d79c-46b4-aa3b-7190ae6061c2","cell_type":"code","source":"# How many studies have the FULL label set vs. report-only?\nhas_all_labels = train[LABEL_COLS].notna().all(axis=1)\nprint(f\"Studies with the full 12-label set : {has_all_labels.sum()}  ({has_all_labels.mean():.1%})\")\nprint(f\"Studies with report text            : {train['Report'].notna().mean():.1%}\")\n\n# Don't stop at \"all 12 or nothing\" -- check per-column availability too.\n# It's possible some studies have a handful of labels filled without having all 12,\n# which the strict .all(axis=1) filter above would silently throw away.\nper_col_available = train[LABEL_COLS].notna().sum().sort_values(ascending=False)\nprint(\"\\nPer-label non-null counts (may exceed the full-set count above):\")\nprint(per_col_available)\nprint(f\"\\nStudies with AT LEAST ONE label filled in: {train[LABEL_COLS].notna().any(axis=1).sum()}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:32.571183Z","iopub.execute_input":"2026-08-05T18:25:32.571522Z","iopub.status.idle":"2026-08-05T18:25:32.586305Z","shell.execute_reply.started":"2026-08-05T18:25:32.571486Z","shell.execute_reply":"2026-08-05T18:25:32.585445Z"}},"outputs":[],"execution_count":null},{"id":"3fda8831-2780-435c-8fbb-ce58e1d35cca","cell_type":"markdown","source":"This ratio is the whole game. If labeled studies are a small slice, the report-derived pseudo-labels below aren't a nice-to-have, they're most of your training signal.\n\nAlso worth noting on a first run: the fully-labeled subset came out at **n=58** — that's tiny, every percentage in section 2 below is essentially a coin-flip-sized sample per label, treat it as a rough shape rather than the true population prevalence. If the per-column check above shows meaningfully more non-null values than 58, that means partial labels exist and it's worth building the training set label-by-label (using `.dropna(subset=[label])` per label) instead of requiring all 12 at once — more usable rows per label that way.","metadata":{}},{"id":"a29ba0d1-216e-49fd-98cd-c7f210a8c9b2","cell_type":"markdown","source":"## 2. Label prevalence (labeled subset only)","metadata":{}},{"id":"fc361a6c-b8ad-496b-8b17-fdb217cd8ef0","cell_type":"code","source":"labeled = train[has_all_labels].copy()\n\nprevalence = labeled[LABEL_COLS].mean().sort_values(ascending=False)\n\nplt.figure(figsize=(9, 5))\nsns.barplot(x=prevalence.values, y=prevalence.index, palette=\"mako\")\nplt.xlabel(\"Positive rate\")\nplt.title(f\"Label prevalence on the labeled subset (n={len(labeled)})\")\nfor i, v in enumerate(prevalence.values):\n    plt.text(v + 0.005, i, f\"{v:.1%}\", va=\"center\", fontsize=9)\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:32.90493Z","iopub.execute_input":"2026-08-05T18:25:32.905772Z","iopub.status.idle":"2026-08-05T18:25:33.375725Z","shell.execute_reply.started":"2026-08-05T18:25:32.905732Z","shell.execute_reply":"2026-08-05T18:25:33.374883Z"}},"outputs":[],"execution_count":null},{"id":"eb08d78c-578c-4416-8896-ba0ab5cf249a","cell_type":"markdown","source":"Don't assume which labels are rare before looking at this — I guessed wrong the first time I wrote this (expected Fracture/Contusion/Baker's to be the rare ones; on a first run they weren't, MCL came out lowest instead). With n this small the ranking could shuffle on a different random subset too, so:\n\n- Whatever ends up lowest here still needs class weighting / oversampling in training, or the model coasts on easy negatives for AUC.\n- Don't treat this ranking as fixed truth about the real (much larger) unlabeled population — it's one small sample.","metadata":{}},{"id":"2a942bbc-b901-4097-9417-116124f32bec","cell_type":"code","source":"corr = labeled[LABEL_COLS].corr()\nplt.figure(figsize=(9, 7))\nsns.heatmap(corr, annot=True, fmt=\".2f\", cmap=\"coolwarm\", center=0, square=True)\nplt.title(\"Label co-occurrence (correlation)\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:34.030938Z","iopub.execute_input":"2026-08-05T18:25:34.031586Z","iopub.status.idle":"2026-08-05T18:25:34.571411Z","shell.execute_reply.started":"2026-08-05T18:25:34.031557Z","shell.execute_reply":"2026-08-05T18:25:34.570523Z"}},"outputs":[],"execution_count":null},{"id":"1b0e63aa-37d5-4263-86cf-f3a151c77613","cell_type":"markdown","source":"Worth checking: do the three OA compartments (Medial/Lateral/PF) correlate with each other, and how strongly? Clinically they often co-occur in advanced disease, but on n=58 don't expect a clean signal — read the actual numbers in the heatmap rather than assuming \"strong\" or \"weak\" going in.","metadata":{}},{"id":"4a2bfa62-675d-462e-9ecb-f763962a3cf5","cell_type":"markdown","source":"## 3. Series-level structure — what are we actually feeding the model","metadata":{}},{"id":"1b42361b-d018-408f-850d-1cd0c85e7cbb","cell_type":"code","source":"series_per_study = train_series.groupby(\"StudyInstanceUID\").size()\n\nfig, axes = plt.subplots(1, 3, figsize=(15, 4))\n\nsns.histplot(series_per_study, bins=20, ax=axes[0], color=\"#4c72b0\")\naxes[0].set_title(\"Series per study\")\naxes[0].set_xlabel(\"# series\")\n\nsns.countplot(data=train_series, x=\"Anatomical_Plane\", ax=axes[1],\n              order=train_series[\"Anatomical_Plane\"].value_counts().index, color=\"#55a868\")\naxes[1].set_title(\"Anatomical plane\")\n\ncombo = train_series.groupby([\"Fluid_Sensitive\", \"Fat_Suppression\"]).size().reset_index(name=\"count\")\ncombo[\"combo\"] = combo.apply(lambda r: f\"FS={r.Fluid_Sensitive} FatSat={r.Fat_Suppression}\", axis=1)\nsns.barplot(data=combo, x=\"combo\", y=\"count\", ax=axes[2], color=\"#c44e52\")\naxes[2].set_title(\"Fluid-sensitive x Fat-sat combos\")\naxes[2].tick_params(axis=\"x\", rotation=30)\n\nplt.tight_layout()\nplt.show()\n\nprint(series_per_study.describe())\nprint(\"\\nFluid_Sensitive x Fat_Suppression combo counts (check this before assuming anything):\")\nprint(combo[[\"Fluid_Sensitive\", \"Fat_Suppression\", \"count\"]])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:35.280072Z","iopub.execute_input":"2026-08-05T18:25:35.280793Z","iopub.status.idle":"2026-08-05T18:25:35.765403Z","shell.execute_reply.started":"2026-08-05T18:25:35.280763Z","shell.execute_reply":"2026-08-05T18:25:35.764513Z"}},"outputs":[],"execution_count":null},{"id":"994cfc3e-bde7-4331-9d6b-2e4e5fc18c0a","cell_type":"markdown","source":"Why this matters for the model, not just curiosity: the dataset card is basically handing you a clinical prior for free.\n\n- Meniscus / ligament tears read best on fluid-sensitive sagittal/coronal sequences.\n- Effusion, contusion, and fracture pop on fluid-sensitive + fat-suppressed (STIR-like) sequences.\n- OA (cartilage loss, joint space narrowing) is more of a structural/T1-ish read, less dependent on fluid sensitivity.\n\nInstead of pooling all series with equal weight into a study-level prediction, gate the series-level attention using `Fluid_Sensitive` / `Fat_Suppression` / `Anatomical_Plane` as auxiliary features.\n\nOne thing to double check rather than assume: on a first run, only two combos showed up (FS=0/FatSat=0 and FS=1/FatSat=1) — the two flags moved together perfectly, no FS=1/FatSat=0 or FS=0/FatSat=1 cases. That's not typical of every real knee protocol, so worth confirming with the printed counts above rather than trusting the bar chart shape alone — it's possible these two fields are simplified/derived rather than fully independent in this dataset.","metadata":{}},{"id":"ce4ea338-3e9d-4598-b44a-6cb60ffd71fc","cell_type":"markdown","source":"## 4. Reports: languages, and can a keyword labeler actually recover the ground truth?","metadata":{}},{"id":"5a89fd71-b765-4cb8-8352-6a47bc7525b3","cell_type":"code","source":"# Multiple reporting sites -> multiple languages, per the dataset card.\n# Cheap language signal without extra dependencies: character set + a handful of\n# language-specific stopwords/markers. Good enough to see the split, not meant to be exact.\ndef rough_lang_guess(text):\n    if not isinstance(text, str) or len(text) < 5:\n        return \"unknown\"\n    t = text.lower()\n    markers = {\n        \"en\": [\" the \", \" and \", \" with \", \" no \", \" findings\"],\n        \"es\": [\" el \", \" la \", \" con \", \" sin \", \" hallazgos\"],\n        \"fr\": [\" le \", \" la \", \" avec \", \" sans \", \" pas de\"],\n        \"de\": [\" der \", \" die \", \" das \", \" mit \", \" kein \"],\n        \"pt\": [\" o \", \" a \", \" com \", \" sem \", \" achados\"],\n    }\n    scores = {lang: sum(t.count(m) for m in ms) for lang, ms in markers.items()}\n    best = max(scores, key=scores.get)\n    return best if scores[best] > 0 else \"other/unknown\"\n\nsample = train[\"Report\"].dropna().sample(min(3000, train[\"Report\"].notna().sum()), random_state=0)\nlang_counts = sample.apply(rough_lang_guess).value_counts(normalize=True)\nlang_counts.plot(kind=\"bar\", color=\"#8172b2\", figsize=(7,4))\nplt.title(\"Rough language mix of a report sample (heuristic, not a real detector)\")\nplt.ylabel(\"share\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:38.091026Z","iopub.execute_input":"2026-08-05T18:25:38.091835Z","iopub.status.idle":"2026-08-05T18:25:38.470868Z","shell.execute_reply.started":"2026-08-05T18:25:38.091804Z","shell.execute_reply":"2026-08-05T18:25:38.469681Z"}},"outputs":[],"execution_count":null},{"id":"5c501847-afd3-43d8-ab56-5778822f08df","cell_type":"markdown","source":"If `other/unknown` + non-English markers together add up to a large share (they did on a first run — well over half the sample wasn't confidently English), that's the single most important number in this notebook: **a pure English keyword list is structurally capped, no amount of tuning the phrase list fixes a language mismatch.** Confirm the language mix of the labeled n=58 subset specifically too (not just the full sample) — the precision/recall table below is only trustworthy for the population it was measured on.","metadata":{}},{"id":"00713725-0395-4847-999c-b2e02ba4f476","cell_type":"code","source":"# A minimal, deliberately simple keyword labeler. The point isn't to be clever,\n# it's to get a *measured* baseline precision/recall so pseudo-labeling decisions\n# are based on numbers instead of assuming report text = ground truth.\nKEYWORDS = {\n    \"ACL\":               [\"acl tear\", \"acl rupture\", \"anterior cruciate ligament tear\",\n                           \"acl injury\", \"torn acl\"],\n    \"MCL\":               [\"mcl tear\", \"mcl sprain\", \"medial collateral ligament\"],\n    \"Medial Meniscus\":   [\"medial meniscus tear\", \"medial meniscal tear\", \"medial meniscus injury\"],\n    \"Lateral Meniscus\":  [\"lateral meniscus tear\", \"lateral meniscal tear\", \"lateral meniscus injury\"],\n    \"Medial OA\":         [\"medial compartment osteoarthritis\", \"medial osteoarthritis\",\n                           \"medial joint space narrowing\"],\n    \"Lateral OA\":        [\"lateral compartment osteoarthritis\", \"lateral osteoarthritis\"],\n    \"PF OA\":             [\"patellofemoral osteoarthritis\", \"patellofemoral compartment osteoarthritis\"],\n    \"Effusion\":          [\"joint effusion\", \"knee effusion\", \"suprapatellar effusion\"],\n    \"Synovitis\":         [\"synovitis\", \"synovial thickening\", \"synovial proliferation\"],\n    \"Baker's\":           [\"baker's cyst\", \"bakers cyst\", \"popliteal cyst\"],\n    \"Contusion\":         [\"bone contusion\", \"bone bruise\", \"osseous contusion\", \"marrow edema\"],\n    \"Fracture\":          [\"fracture\", \"cortical break\"],\n}\nNEGATION_WINDOW = [\"no \", \"without \", \"negative for \", \"no evidence of \"]\n\ndef keyword_predict(report, label):\n    if not isinstance(report, str):\n        return 0\n    text = report.lower()\n    for kw in KEYWORDS[label]:\n        idx = text.find(kw)\n        if idx == -1:\n            continue\n        window = text[max(0, idx - 25): idx]\n        if any(neg in window for neg in NEGATION_WINDOW):\n            continue\n        return 1\n    return 0\n\nfrom sklearn.metrics import precision_score, recall_score, f1_score\n\nrows = []\neval_df = labeled.dropna(subset=[\"Report\"])\nfor label in LABEL_COLS:\n    y_true = eval_df[label].astype(int)\n    y_pred = eval_df[\"Report\"].apply(lambda r: keyword_predict(r, label))\n    rows.append({\n        \"label\": label,\n        \"precision\": precision_score(y_true, y_pred, zero_division=0),\n        \"recall\": recall_score(y_true, y_pred, zero_division=0),\n        \"f1\": f1_score(y_true, y_pred, zero_division=0),\n        \"n_positive\": int(y_true.sum()),\n        \"n_predicted_positive\": int(y_pred.sum()),\n    })\n\nreport_quality = pd.DataFrame(rows).sort_values(\"f1\", ascending=False)\nreport_quality\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:40.70754Z","iopub.execute_input":"2026-08-05T18:25:40.708344Z","iopub.status.idle":"2026-08-05T18:25:40.841556Z","shell.execute_reply.started":"2026-08-05T18:25:40.708312Z","shell.execute_reply":"2026-08-05T18:25:40.840676Z"}},"outputs":[],"execution_count":null},{"id":"acd7e38f-addb-4151-bad6-51d469619abe","cell_type":"code","source":"plt.figure(figsize=(9, 5))\nmelted = report_quality.melt(id_vars=\"label\", value_vars=[\"precision\", \"recall\", \"f1\"])\nsns.barplot(data=melted, x=\"value\", y=\"label\", hue=\"variable\")\nplt.title(\"Keyword labeler vs. ground truth, per label\")\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:42.813796Z","iopub.execute_input":"2026-08-05T18:25:42.814125Z","iopub.status.idle":"2026-08-05T18:25:43.17785Z","shell.execute_reply.started":"2026-08-05T18:25:42.814098Z","shell.execute_reply":"2026-08-05T18:25:43.177064Z"}},"outputs":[],"execution_count":null},{"id":"d63212c3-7271-4034-a6fa-2934ac00c8b8","cell_type":"markdown","source":"Read the actual table before pseudo-labeling anything — don't assume \"reports = free labels\" just because reports exist:\n\n- **A label with `n_predicted_positive` near 0** means the keyword list didn't fire almost at all for that label on this sample. Zero precision AND zero recall together is a sign of a complete miss, not just a weak one — usually because the exact English phrases in the list don't appear (wrong language, or clinicians phrasing it completely differently, e.g. \"joint space narrowing\" instead of the word \"osteoarthritis\").\n- **High precision, low recall** → safe to use as one-sided pseudo-*positive* labeling only (trust the hits, don't assume a miss means negative — recall this low means plenty of true positives get missed and would be mislabeled negative).\n- **Both low** → shortlist that label for a proper LLM-based / multilingual extractor before using it for anything.\n\nThis table is the actual deliverable of this section: a *per-label* trust score for weak supervision, measured, not assumed.","metadata":{}},{"id":"5d07ed3e-db74-4f96-a4f9-c9e260b584a3","cell_type":"markdown","source":"## 5. What a study actually looks like — slice grid across planes","metadata":{}},{"id":"edac918e-6f99-4d47-a7bd-89f2bcc1c329","cell_type":"code","source":"def load_series_slices(study_uid, series_uid, max_slices=9):\n    series_dir = f\"{DATA_DIR}/train_series/{study_uid}/{series_uid}\"\n    files = sorted(glob.glob(f\"{series_dir}/*.dcm\"))\n    if not files:\n        return []\n    step = max(1, len(files) // max_slices)\n    picked = files[::step][:max_slices]\n    slices = []\n    for fp in picked:\n        try:\n            ds = pydicom.dcmread(fp)\n            arr = ds.pixel_array.astype(np.float32)\n            arr = (arr - arr.min()) / (arr.max() - arr.min() + 1e-6)\n            slices.append(arr)\n        except Exception as e:\n            print(f\"skipped {fp}: {e}\")\n    return slices\n\n# Pick one study that has multiple series with different planes, for a good comparison\nexample_study = train_series[\"StudyInstanceUID\"].value_counts().index[0]\nexample_series = train_series[train_series[\"StudyInstanceUID\"] == example_study]\nprint(example_series[[\"SeriesInstanceUID\", \"Anatomical_Plane\", \"Fluid_Sensitive\", \"Fat_Suppression\"]])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:45.920771Z","iopub.execute_input":"2026-08-05T18:25:45.921143Z","iopub.status.idle":"2026-08-05T18:25:45.939801Z","shell.execute_reply.started":"2026-08-05T18:25:45.921114Z","shell.execute_reply":"2026-08-05T18:25:45.93887Z"}},"outputs":[],"execution_count":null},{"id":"4a1e5b22-832c-42dd-a5ff-1500aa73af41","cell_type":"code","source":"fig, axes = plt.subplots(1, min(4, len(example_series)), figsize=(16, 4))\nif len(example_series) == 1:\n    axes = [axes]\n\nfor ax, (_, row) in zip(axes, example_series.iloc[:4].iterrows()):\n    slices = load_series_slices(example_study, row[\"SeriesInstanceUID\"], max_slices=1)\n    if slices:\n        ax.imshow(slices[0], cmap=\"gray\")\n    ax.set_title(f\"{row['Anatomical_Plane']}\\nFS={row['Fluid_Sensitive']} FatSat={row['Fat_Suppression']}\",\n                 fontsize=9)\n    ax.axis(\"off\")\n\nplt.tight_layout()\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-08-05T18:25:47.721148Z","iopub.execute_input":"2026-08-05T18:25:47.721494Z","iopub.status.idle":"2026-08-05T18:25:48.448423Z","shell.execute_reply.started":"2026-08-05T18:25:47.721442Z","shell.execute_reply":"2026-08-05T18:25:48.446516Z"}},"outputs":[],"execution_count":null},{"id":"61870d43-527a-4b93-b60c-fb314abd8178","cell_type":"markdown","source":"If any of those panels came back blank with a decode error in the printed log above, that's the compressed-transfer-syntax issue mentioned at the top — double check the `pylibjpeg`/`gdcm` install actually took before assuming your data loader is broken.","metadata":{}},{"id":"5a39a15e-da2e-4d58-add7-76f143a82ba6","cell_type":"markdown","source":"## 6. Takeaways\n\n- Report-only studies vastly outnumber fully labeled ones — pseudo-labeling isn't optional here, it's most of the usable signal, but section 4's table shows it's not equally trustworthy across all 12 labels.\n- Check per-column label availability, not just the strict \"all 12 present\" filter — there may be more usable rows hiding per-label.\n- `Fluid_Sensitive` / `Fat_Suppression` / `Anatomical_Plane` are a free clinical prior for weighting series in aggregation — but verify the combo counts before assuming they behave the way a textbook knee protocol would.\n- Language mix matters more than phrase-list cleverness. Check it before investing time in expanding an English-only keyword list.\n- Don't skip the DICOM decoder install. It's a five-second `pip install` that saves an afternoon of \"why is my dataloader silently dropping half the studies.\"\n\nIf this was useful, an upvote helps it reach more people starting out on this comp — and if you spot a label where the keyword list is obviously missing something (Synovitis and Contusion phrasing especially, and non-English OA terms like \"artrosis\" / \"arthrose\" / \"Gonarthrose\"), drop it in the comments, happy to fold it into a v2 pass on the labeler.\n","metadata":{}}]}