{"cells":[{"cell_type":"markdown","metadata":{},"source":"# Only 58 labels? Mine the reports.\n\nA thing that is easy to miss about this competition: `train.csv` has 4,407 studies but **only 58 of them\ncarry ground-truth labels**. The other 4,349 rows have an empty label block and a free-text **radiology\nreport** instead. The test studies have neither, so at inference you predict from images alone.\n\nThat makes this a **weakly-supervised** problem. The reports are the training signal, and they are\nmultilingual (Dutch, Spanish, English, and more). This notebook does three things:\n\n1. shows the label / report structure and the language mix,\n2. gives a small, transparent, multilingual **rule-based labeler** that turns a report into the 12 labels,\n3. validates it on the 58 gold studies (macro-AUC), and writes pseudo-labels you can train an image model on.\n\nNothing here is heavy. It runs in under a minute on CPU. Upvote if it saves you some time, and tell me in\nthe comments where the rules miss so we can sharpen them.\n"},{"cell_type":"markdown","metadata":{},"source":"## Load the data"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"import pandas as pd, numpy as np, re, glob\n# The competition data can mount either flat (/kaggle/input/<slug>/) or nested\n# (/kaggle/input/competitions/<slug>/), so find train.csv wherever it landed.\n_cands = glob.glob(\"/kaggle/input/**/train.csv\", recursive=True)\n_cands = [c for c in _cands if \"rsna-knee-abnormality-detection\" in c] or _cands\nassert _cands, \"train.csv not found - attach the competition data to this notebook\"\ntrain = pd.read_csv(_cands[0])\nprint(\"loaded:\", _cands[0])\nLABELS = [\"ACL\",\"MCL\",\"Medial Meniscus\",\"Lateral Meniscus\",\"Medial OA\",\"Lateral OA\",\n          \"PF OA\",\"Effusion\",\"Synovitis\",\"Baker's\",\"Contusion\",\"Fracture\"]\ngold_mask = train[LABELS].notna().all(axis=1)\nprint(\"studies:\", len(train))\nprint(\"with all 12 labels (gold):\", int(gold_mask.sum()))\nprint(\"report only:\", int((~gold_mask).sum()))\nprint(\"reports present:\", int(train[\"Report\"].notna().sum()))"},{"cell_type":"markdown","metadata":{},"source":"## Label prevalence (on the 58 gold) and report languages"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"gold = train[gold_mask]\nprev = gold[LABELS].mean().sort_values(ascending=False)\nprint(\"positive rate on gold:\")\nfor k,v in prev.items(): print(f\"  {k:18s} {v:.2f}\")\n\ndef guess_lang(t):\n    t=str(t).lower()\n    if any(w in t for w in [\"rodilla\",\"menisco\",\"rotura\",\"derrame\",\"técnica\",\"tecnica\"]): return \"es\"\n    if any(w in t for w in [\"knie\",\"meniscus\",\"scheur\",\"bevindingen\",\"geen\"]): return \"nl\"\n    if any(w in t for w in [\"knee\",\"tear\",\"effusion\",\"findings\",\"meniscal\"]): return \"en\"\n    return \"other/mixed\"\ntrain[\"lang\"]=train[\"Report\"].map(guess_lang)\nprint(\"\\nreport language (rough):\", train[\"lang\"].value_counts().to_dict())\nprint(\"median report length (chars):\", int(train[\"Report\"].str.len().median()))"},{"cell_type":"markdown","metadata":{},"source":"## A transparent multilingual labeler\n\nThe idea is simple: for each finding, look for the anatomy term near an abnormality term, and reject it if a\nnegation cue sits just before it (`no tear`, `sin rotura`, `geen scheur`, `intact`, `preserved`, `normal`).\nSingle-concept findings (effusion, Baker's cyst, fracture, synovitis) fire on presence alone. It is\ndeliberately readable so you can extend it - the patterns cover English, Spanish and Dutch first."},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"\"\"\"\nMultilingual rule-based report labeler for RSNA Knee Abnormality Detection.\n\nTurns a free-text knee-MRI radiology report (English / Spanish / Dutch / others)\ninto the 12 competition labels. This is the *baseline* labeler: transparent regex\npatterns with per-finding abnormality cues and a negation guard. It is meant to\n(a) prove the report-mining thesis on the 58 gold studies and (b) serve as the\npublic notebook baseline. The stronger LLM labeler is a separate, private module.\n\nLabels: ACL, MCL, Medial Meniscus, Lateral Meniscus, Medial OA, Lateral OA,\nPF OA, Effusion, Synovitis, Baker's, Contusion, Fracture.\n\"\"\"\nimport re\n\nLABELS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n          \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n          \"Contusion\", \"Fracture\"]\n\n# Negation cues (multilingual). If one of these appears within NEG_WINDOW chars\n# BEFORE a finding hit (same sentence-ish), we treat the finding as negated.\nNEG = [\n    # english\n    r\"\\bno\\b\", r\"\\bnot\\b\", r\"\\bwithout\\b\", r\"\\bintact\\b\", r\"\\bnormal\\b\",\n    r\"\\bpreserved\\b\", r\"\\bunremarkable\\b\", r\"\\bno evidence\\b\", r\"\\brules? out\\b\",\n    # spanish\n    r\"\\bsin\\b\", r\"\\bno\\b\", r\"\\bnormal(es)?\\b\", r\"\\bconservad\", r\"\\b[ií]ntegr\",\n    r\"\\bpreservad\", r\"\\bausencia\\b\", r\"\\bdescarta\",\n    # dutch\n    r\"\\bgeen\\b\", r\"\\bzonder\\b\", r\"\\bintact\\b\", r\"\\bnormaal\\b\", r\"\\bnormale\\b\",\n    r\"\\bongestoord\\b\", r\"\\bniet\\b\",\n]\nNEG_RE = re.compile(\"|\".join(NEG), re.I)\nNEG_WINDOW = 45          # chars before the hit to scan for a negation cue\n\n# Per-label patterns: (anatomy_terms, abnormality_terms).\n# A label fires if an anatomy term co-occurs (within SENT window) with an\n# abnormality term and is NOT negated. Some labels (effusion, baker, fracture,\n# synovitis) are single-concept: presence term == the finding itself.\nTEAR = r\"(tear|torn|rupture|ruptur|disrupt|rotura|roto|rota|desgarro|desinserci|scheur|ruptuur|letsel)\"\nDEGEN = r\"(degenerat|degenerativ|mucoid|fissur|fisura|maceraci)\"\n\nPATT = {\n    \"ACL\": (r\"(acl|anterior cruciate|cruzado anterior|\\blca\\b|voorste kruisband|\\bvkb\\b)\", TEAR),\n    \"MCL\": (r\"(mcl|medial collateral|colateral medial|colateral interno|mediale collaterale|\\blcm\\b)\",\n            r\"(tear|torn|rupture|ruptur|sprain|rotura|roto|esguince|scheur|ruptuur|letsel)\"),\n    \"Medial Meniscus\": (r\"(medial meniscus|menisco medial|menisco interno|mediale meniscus|binnenmeniscus)\",\n                        TEAR + \"|\" + DEGEN),\n    \"Lateral Meniscus\": (r\"(lateral meniscus|menisco lateral|menisco externo|laterale meniscus|buitenmeniscus)\",\n                         TEAR + \"|\" + DEGEN),\n    \"Medial OA\": (r\"(medial compartment|femorotibial medial|compartiment.{0,15}medial|mediale compartiment|femorotibiaal mediaal)\",\n                  r\"(osteoarthrit|arthros|artros|artrose|chondral loss|condral|cartilage loss|gonartros|osteofit|osteophyt)\"),\n    \"Lateral OA\": (r\"(lateral compartment|femorotibial lateral|compartiment.{0,15}lateral|laterale compartiment|femorotibiaal lateraal)\",\n                   r\"(osteoarthrit|arthros|artros|artrose|chondral loss|condral|cartilage loss|gonartros|osteofit|osteophyt)\"),\n    \"PF OA\": (r\"(patellofemoral|patelofemoral|femoropatelar|femoropatellair|retropatell|patellar facet|tr[oó]clea)\",\n              r\"(osteoarthrit|arthros|artros|artrose|chondral|condral|cartilage|osteofit|osteophyt|condropat|chondropath)\"),\n    # single-concept findings\n    \"Effusion\": (r\"(effusion|derrame|vocht|erguss|epanchement|hydrops)\", r\".\"),\n    \"Synovitis\": (r\"(synovitis|sinovitis|synovit|synoviale|hoffitis|plica)\", r\".\"),\n    \"Baker's\": (r\"(baker|popliteal cyst|quiste popl[ií]teo|bakercyste|poplitea?le? cyste)\", r\".\"),\n    \"Contusion\": (r\"(bone contusion|bone bruise|bone marrow ede|marrow ede|contusi[oó]n [oó]sea|edema [oó]seo|botcontusie|beenmergoedeem|beenmerg[o]?edeem)\", r\".\"),\n    \"Fracture\": (r\"(fracture|fractura|fractuur|avulsion|avulsi[oó]n)\", r\".\"),\n}\nCOMPILED = {k: (re.compile(a, re.I), re.compile(b, re.I)) for k, (a, b) in PATT.items()}\n\n\ndef _negated(text, pos):\n    \"\"\"Is there a negation cue in the window just before position `pos`?\"\"\"\n    start = max(0, pos - NEG_WINDOW)\n    return NEG_RE.search(text[start:pos]) is not None\n\n\ndef label_report(report: str) -> dict:\n    \"\"\"Return {label: 0/1} for one report.\"\"\"\n    t = str(report)\n    out = {}\n    for lab, (anat_re, abn_re) in COMPILED.items():\n        fired = 0\n        for m in anat_re.finditer(t):\n            a, b = m.start(), m.end()\n            # window around the anatomy mention to look for abnormality term\n            window = t[max(0, a - 60): b + 90]\n            if abn_re.pattern == \".\" :          # single-concept finding\n                if not _negated(t, a):\n                    fired = 1\n                    break\n            else:\n                am = abn_re.search(window)\n                if am and not _negated(t, a) and not _negated(window, am.start()):\n                    fired = 1\n                    break\n        out[lab] = fired\n    return out\n\n\ndef label_frame(df, report_col=\"Report\"):\n    import pandas as pd\n    rows = [label_report(r) for r in df[report_col].fillna(\"\")]\n    return pd.DataFrame(rows, index=df.index)[LABELS]"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"def _auc(y, s):\n    pairs = sorted(zip(s, y), key=lambda t: t[0]); n=len(pairs); ranks=[0.0]*n; i=0\n    while i < n:\n        j=i\n        while j<n and pairs[j][0]==pairs[i][0]: j+=1\n        r=(i+1+j)/2.0\n        for k in range(i,j): ranks[k]=r\n        i=j\n    P=sum(y); N=len(y)-P\n    if P==0 or N==0: return float(\"nan\")\n    rp=sum(ranks[k] for k in range(n) if pairs[k][1]==1)\n    return (rp - P*(P+1)/2.0)/(P*N)"},{"cell_type":"markdown","metadata":{},"source":"## Validate on the 58 gold studies"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"pred = label_frame(gold).astype(float)\ny = gold[LABELS].astype(int).values\nprint(f\"{'label':18s} pos%  auc\")\naucs=[]\nfor i,c in enumerate(LABELS):\n    a=_auc(list(y[:,i]), list(pred[c].values)); aucs.append(a)\n    print(f\"{c:18s} {100*y[:,i].mean():4.0f}  {a:.2f}\")\nmacro=np.nanmean(aucs)\nprint(\"-\"*30)\nprint(f\"{'MACRO-AUC':18s}       {macro:.3f}   (random = 0.500)\")"},{"cell_type":"markdown","metadata":{},"source":"The strong labels (MCL, Baker's cyst, ACL) already sit well above chance from plain rules. The weak ones\n(medial vs lateral osteoarthritis, effusion) are where naive keywords struggle: they need compartment\nreasoning and better negation. That is the natural place to bring in a multilingual LLM labeler, which reads\nthe finding in context and returns a calibrated probability instead of a hard 0/1. Left as an exercise here."},{"cell_type":"markdown","metadata":{},"source":"## Write pseudo-labels for the 4,349 report-only studies"},{"cell_type":"code","metadata":{},"execution_count":null,"outputs":[],"source":"allpred = label_frame(train).astype(float)\nout = pd.concat([train[[\"StudyInstanceUID\"]].reset_index(drop=True),\n                 allpred.reset_index(drop=True)], axis=1)\n# keep the 58 gold as their true labels\nfor c in LABELS:\n    out.loc[gold_mask.values, c] = train.loc[gold_mask, c].values\nout.to_csv(\"pseudo_labels_rules.csv\", index=False)\nprint(\"wrote pseudo_labels_rules.csv\", out.shape)\nout.head()"},{"cell_type":"markdown","metadata":{},"source":"## Where to go next\n\n- **Better labeler.** Swap the rules for a multilingual medical LLM and score a probability per finding.\n  Validate the same way, on the 58 gold, and only keep it if the macro-AUC goes up.\n- **Compartment logic.** Medial vs lateral OA is the weakest link. The series table gives you the\n  anatomical plane per series, which helps on the image side too.\n- **Image model.** Train on these pseudo-labels, predict from images at test time (no report available there).\n\nIf the rules miss a finding in your language, drop an example in the comments and I will fold it in.\n"}],"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3"},"language_info":{"name":"python"}},"nbformat":4,"nbformat_minor":5}