{"cells":[{"cell_type":"code","id":"cover-00","metadata":{"_kg_hide-input":true,"_kg_hide-output":false},"execution_count":null,"outputs":[],"source":"from pathlib import Path\nfrom IPython.display import Image, display\n\n\ndef _find_cover(name=\"RSNA_KNEE_1.png\"):\n    \"\"\"Locate the cover image wherever its dataset was mounted.\n\n    A hardcoded mount path guarded by `exists()` is the worst of both worlds: get it\n    wrong and the image simply is not there, with nothing said. Kaggle also does not\n    always mount a dataset at the same depth. Searching the attached inputs costs one\n    directory listing per input and cannot fail quietly. The competition mount is\n    skipped by inspection rather than by name, because it holds hundreds of thousands\n    of files and none of them is this one.\n    \"\"\"\n    base = Path(\"/kaggle/input\")\n    if not base.is_dir():\n        return None\n    for d in sorted(p for p in base.iterdir() if p.is_dir()):\n        if any((d / s).is_dir() for s in (\"train_series\", \"test_series\")):\n            continue\n        hit = next(d.rglob(name), None)\n        if hit is not None:\n            return hit\n    return None\n\n\n_cover = _find_cover()\nif _cover is not None:\n    display(Image(filename=str(_cover)))\n"},{"cell_type":"markdown","id":"md-01","metadata":{},"source":"# Twelve findings from one knee MRI\n\nEach study is a set of MRI series acquired in one session, and the task is to give it\ntwelve probabilities: anterior cruciate and medial collateral ligament injury, medial and\nlateral meniscal tear, osteoarthritis in each of the three compartments, joint effusion,\nsynovitis, Baker's cyst, bone contusion, and fracture.\n\nThe sections are ordered by how the decisions constrain each other: what the score rewards\ndecides how predictions are combined, where the targets come from decides what can be\ntrained, and what the scanner recorded decides what the encoder is shown.\n"},{"cell_type":"markdown","id":"md-02","metadata":{},"source":"> **On the two label sources.** §2 describes two readers for one job: a rule extractor,\n> defined in full below, and a language model reading the same reports, whose output is a\n> public table attached as a dataset. With the table mounted it supplies the targets,\n> without it the extractor does, and every cell after that is unchanged.\n"},{"cell_type":"markdown","id":"md-03","metadata":{},"source":"## 1. What the score rewards\n\nThe score is the unweighted mean of twelve per-label ROC AUCs:\n\n$$\\text{Score} \\;=\\; \\frac{1}{12}\\sum_{i=0}^{11} \\mathrm{AUC}_i .$$\n\nThree consequences follow, and each removes a design choice.\n\n**Only order matters.** $\\mathrm{AUC}_i$ is invariant under any strictly increasing map of\nthe scores for label $i$, so calibration and thresholds are worth nothing. It also fixes\nhow to combine models: averaging raw probabilities lets whichever model is most confident\ndominate, whereas averaging *ranks* combines the only information the metric reads.\n\n**Every label costs the same.** Write $M$ for the mean AUC a good model could reach. A\nlabel left at chance contributes $0.5$ instead of roughly $M$, forfeiting $(M-0.5)/12$ of\nthe final score however well the other eleven do — at $M = 0.85$, $0.029$. Rare findings\ndeserve *more* attention than common ones, because a rare finding is where a model most\neasily ends up at chance.\n\n**Prevalence drifts are survivable, thresholds are not.** AUC is in expectation invariant\nto the positive rate, and the competition states prevalence is not guaranteed to match\nacross the training, public and final sets. One cutoff survives below — the graded targets\nare binarised at their midpoint so that a held-out AUC can be computed at all — and it\ndecides which epoch and configuration are kept, never a submitted score.\n"},{"cell_type":"markdown","id":"md-04","metadata":{},"source":"## 2. Where the targets come from\n\nOnly a small subset of training studies carry the twelve per-condition labels. Every\ntraining study carries the original radiology report, and the data description invites\nderiving labels from it.\n\nThe decisive fact is in the schemas rather than the prose: `train.csv` has a `Report`\ncolumn and `test.csv` does not. Text is available when fitting and absent when predicting.\nThat rules out a fusion model with a text branch — at inference it would have nothing to\nread — and leaves the reports usable only as training targets, as an auxiliary signal\ndropped at inference, or as a weight on how confidently a study could be read. This\nnotebook takes the first and the third: a multilingual rule extractor reads each report\nclause by clause, deciding for each finding whether the clause asserts, negates or hedges\nit, and emits a score with a confidence. The confidence becomes a sample weight, so a study\nwhose report says nothing about synovitis pulls on that head far less than one that names it.\n\n**Two readers.** A lexicon matches morphology, so its failure mode is silence rather than\nerror, and silence is measurable without ground truth: for each (report, finding) pair, ask\nonly whether anything matched at all. That rate needs no annotations, so it runs on every\nstudy, and tagged by language it says *where* the vocabulary is thin. Here the misses are\nconcentrated — one language is covered far better than the other eight, and the gap falls on\nfindings a knee report almost always comments on. Enumerating morphology for nine languages\nis the wrong instrument for that; reading the sentence is the right one, and a language model\nreads it, asked for the same twelve findings in the same graded form. The annotated studies\ndecide between the readers, paired on the same studies, and the difference is large and\none-sided in the direction the coverage rate predicts. So the pipeline prefers a mounted\ntable of model-read labels when present and runs the lexicon when not; both emit the same\ncolumns, and a partial table falls back per study rather than per run.\n\nTwo details matter more than either reader's internals.\n\n**Reports are graded, annotations are thresholded.** The reporting radiologist and the\nannotator do not share a threshold. A report saying *small joint effusion* may sit against\na negative annotation, because the annotator marked only effusions they judged significant.\nA rule of the form *term present $\\Rightarrow$ positive* is wrong by construction; grading\nthe mention — trace, unqualified, marked — is right and costs nothing, because §1\nestablished that only order is read.\n\n**Derived labels are not independent across studies.** A report shared verbatim by several\nstudies yields one target vector for all of them, which must be respected when splitting;\n§7 does.\n"},{"cell_type":"markdown","id":"md-05","metadata":{},"source":"### Reading a report in nine languages\n\n**No language is identified.** Every cue lexicon carries all nine languages at once and\neach clause is tested against the union. Routing first means committing to a guess before\nany evidence is read, and the cheap guess — substring tests, `'the '` for English, `'la '`\nfor French — fails badly, because `la` is as common in Spanish as in French and whichever\ntest runs first swallows both. Pooling costs little, since Greek and Cyrillic cues cannot\ncollide with Latin-script ones and the Latin-script vocabularies of interest are close\nenough that a shared cue is usually right. The price is paid in coverage instead.\n\n**Normalise, then segment, then scope.** Case, diacritics and separators are folded first,\nwhich also repairs a codepoint problem: many Greek reports spell mu with MICRO SIGN U+00B5\nrather than U+03BC, and NFKD maps one onto the other. Text is then split into clauses, with\na heading line attached to the value beneath it, because a report reading `Fractures :`\nthen `Aucune.` states one thing across two lines and any method that splits them reads a\nnegation as a positive.\n\n**Assertion, negation, hedge.** Negation is not an edge case: for several findings most\nmentions are negative, since a report lists what was checked and found intact. Explicit\nnormality counts as negation — *ligamentos cruzados y colaterales dentro de límites\nnormales* is evidence of absence, not absence of evidence — except where a tear or a high\ngrade is named in the same breath.\n\n**Stems behind phrases.** Four targets need an anatomy word and a pathology word together.\nA lexicon of complete phrases carries most matches but cannot survive morphology: Turkish\nsuffixes possessives onto the noun, Croatian and Greek decline it. Where the phrase fails,\na second pass matches a stem and requires a side qualifier within a *character* window,\nwhich handles inflection without enumerating it, and handles word order — which puts the\nside adjective before the noun in English and after it in Greek.\n\n**Why coverage decides and agreement does not.** A rule that never fires raises no error;\nin a binary extractor it emits a negative, indistinguishable from a confident one, and a\nlexicon complete in English and thin in Greek looks like a corpus where Greek patients have\nfewer findings. The obvious check — agreement with the per-condition annotations — measures\nthe right thing on far too few studies. The Hanley–McNeil standard error of an AUC $A$ with\n$n_p$ positives and $n_n$ negatives,\n\n$$\n\\mathrm{SE}(A)=\\sqrt{\\frac{A(1-A)+(n_p-1)(Q_1-A^{2})+(n_n-1)(Q_2-A^{2})}{n_p\\,n_n}},\n\\qquad\nQ_1=\\frac{A}{2-A},\\quad Q_2=\\frac{2A^{2}}{1+A},\n$$\n\nat $A\\approx0.8$ and a handful of positives among a few dozen studies lands near $0.09$ — a\n95% interval of roughly $\\pm0.17$ — far wider than the differences between two lexicons it would be asked to\nseparate. The silence rate has no such limit, so it is what points at the missing\nvocabulary and what changes to the lexicon are judged on.\n"},{"cell_type":"code","id":"code-06","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"from __future__ import annotations\n\nimport re\nimport unicodedata\n\nTARGETS = [\n    \"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\",\n    \"Medial OA\", \"Lateral OA\", \"PF OA\", \"Effusion\",\n    \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\",\n]\n\n\n# Turkish dotted/dotless i must be folded before casefolding, otherwise \"İZLENMEZ\"\n# and \"izlenmez\" diverge. ß and the Croatian/Serbian d-with-stroke likewise.\n_PRE = str.maketrans({\n    \"ı\": \"i\", \"İ\": \"i\", \"I\": \"i\", \"ß\": \"ss\", \"đ\": \"d\", \"Đ\": \"d\",\n    \"ø\": \"o\", \"Ø\": \"o\", \"æ\": \"ae\", \"Æ\": \"ae\",\n})\n\n\ndef normalize(text: str) -> str:\n    \"\"\"Fold case, diacritics and separators; keep Greek and Cyrillic letters.\n\n    NFKD decomposition strips Latin accents and Greek tonos alike (ά -> α), which is what\n    we want: reports are inconsistent about accents. It also maps the MICRO SIGN U+00B5\n    to a real mu, which matters because most Greek reports here use the wrong codepoint.\n    \"\"\"\n    if not isinstance(text, str):\n        return \"\"\n    text = text.translate(_PRE).lower()\n    text = unicodedata.normalize(\"NFKD\", text)\n    text = \"\".join(ch for ch in text if not unicodedata.combining(ch))\n    text = text.replace(\"­\", \"\")                    # soft hyphen\n    text = re.sub(r\"[_\\-/\\\\]+\", \" \", text)\n    text = re.sub(r\"[ \\t]+\", \" \", text)\n    return text\n\n\n_SENT_SPLIT = re.compile(r\"(?<=[.;!?])\\s+|\\n+\")\n\n\ndef clauses(text: str):\n    \"\"\"Split into clauses, then attach `header:` lines to the value that follows.\n\n    A report line reading `Fractures :` followed by `Aucune.` is one statement. Splitting\n    on punctuation alone separates the anatomy from its negation and flips the label.\n    \"\"\"\n    norm = normalize(text)\n    raw = [c.strip() for c in _SENT_SPLIT.split(norm) if c and c.strip()]\n\n    merged = []\n    for i, c in enumerate(raw):\n        # A fragment ending in a colon is a heading for the next fragment. Structured\n        # English reports write long ones - \"lateral compartment (meniscus, collateral\n        # ligament complex, cartilage):\" is eight words - so the cap is generous.\n        #\n        # A merged heading must NOT also stand alone. On its own it carries the anatomy\n        # word with no negation in scope, so `Fractures :` / `Aucune.` asserted a fracture\n        # off the heading while the joined clause correctly read the denial. The joined\n        # clause is a superset of the heading, so nothing is lost by dropping it; a\n        # heading with no value beneath it is not merged and still stands.\n        if c.endswith(\":\") and len(c.split()) <= 14 and i + 1 < len(raw):\n            merged.append(c + \" \" + raw[i + 1])\n        else:\n            merged.append(c)\n    # Comma-separated enumerations inside a long clause hide separate assertions.\n    out = []\n    for c in merged:\n        out.append(c)\n        if len(c.split()) > 25:\n            out.extend(p.strip() for p in c.split(\",\") if len(p.split()) > 2)\n    return out\n\n\ndef _rx(*alts: str) -> re.Pattern:\n    return re.compile(\"|\".join(alts))\n"},{"cell_type":"code","id":"code-07","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"NEGATION = _rx(\n    # en\n    r\"\\bno\\b\", r\"\\bnot\\b\", r\"\\bwithout\\b\", r\"\\bnegative for\\b\", r\"\\babsence\\b\",\n    r\"\\bno evidence\\b\", r\"\\bunremarkable\\b\", r\"\\bfree of\\b\", r\"\\bnone\\b\", r\"\\bnil\\b\",\n    # es\n    r\"\\bsin\\b\", r\"\\bno hay\\b\", r\"\\bausencia\\b\", r\"\\bausentes?\\b\",\n    # fr\n    r\"\\bpas de\\b\", r\"\\bsans\\b\", r\"\\baucune?\\b\", r\"\\babsence\\b\",\n    # nl\n    r\"\\bgeen\\b\", r\"\\bzonder\\b\", r\"\\bniet\\b\",\n    # de\n    r\"\\bkeine?\\b\", r\"\\bohne\\b\", r\"\\bnicht\\b\",\n    # tr\n    r\"\\byok\\b\", r\"\\byoktur\\b\", r\"izlenmemekte\", r\"saptanmadi\", r\"\\bdegil\\b\",\n    r\"gozlenmemekte\", r\"mevcut degil\", r\"eslik etmiyor\", r\"\\bizlenmedi\\b\",\n    # hr / sr / bs\n    r\"\\bnema\\b\", r\"\\bbez\\b\", r\"\\bnisu\\b\", r\"\\bnije\\b\",\n    # el (accents already stripped)\n    r\"\\bδεν\\b\", r\"\\bχωρις\\b\", r\"ουδεν\",\n    # bg / ru\n    r\"\\bбез\\b\", r\"\\bне\\b\", r\"липсва\", r\"\\bняма\\b\",\n)\n\nNORMALITY = _rx(\n    r\"\\bnormal\", r\"\\bintact\\b\", r\"\\bpreserved\\b\", r\"\\bwithin normal limits\\b\",\n    r\"limites normales\", r\"\\bconservad\", r\"\\bintegr\", r\"\\bnormales\\b\",\n    r\"\\bdoga(l|ll)\\b\", r\"korunmus\", r\"\\bnormaldir\\b\", r\"olagan\",\n    r\"\\buredn\", r\"\\bocuvan\", r\"\\bodrzan\", r\"\\bintakt\",\n    r\"φυσιολογικ\", r\"ακεραι\",\n    r\"unauffallig\", r\"regelrecht\", r\"\\bintakt\\b\",\n    r\"нормал\", r\"запазен\", r\"съхранен\", r\"\\bбез особености\\b\",\n    r\"\\bgaaf\\b\", r\"\\bnormaal\\b\",\n)\n\nUNCERTAIN = _rx(\n    r\"\\bpossible\\b\", r\"\\bprobable\\b\", r\"\\bsuspicious\\b\", r\"\\bsuspected\\b\",\n    r\"cannot (be )?exclude\", r\"\\bmay\\b\", r\"\\bquestionable\\b\", r\"\\bequivocal\\b\",\n    r\"\\bposible\\b\", r\"sin criterios categoricos\", r\"\\bdudos\",\n    r\"\\bmuhtemel\\b\", r\"\\bolasi\\b\", r\"\\bsupheli\\b\", r\"\\bizlenim\",\n    r\"\\bmoguce\\b\", r\"\\bvjerojatno\\b\", r\"\\bsumnja\\b\",\n    r\"πιθαν\", r\"υποπτ\",\n    r\"\\bmoglich\", r\"\\bverdachtig\", r\"\\bfraglich\", r\"\\bV\\.a\\.\\b\",\n    r\"\\bвъзможно\\b\", r\"\\bвероятно\\b\", r\"суспект\",\n    r\"\\bmogelijk\\b\", r\"\\bverdacht\\b\",\n)\n\n# Pathology vocabulary shared by the paired rules.\nTEAR = _rx(\n    r\"\\btear\", r\"\\btorn\\b\", r\"\\brupture\", r\"\\bdisruption\\b\", r\"discontinuit\",\n    r\"\\bavuls\",\n    r\"\\brotura\\b\", r\"\\broturas\\b\", r\"\\bruptura\", r\"\\bdesgarro\", r\"\\broto\\b\",\n    r\"\\bdechirure\", r\"\\bdechire\",\n    r\"\\bscheur\", r\"\\bruptuur\", r\"gescheurd\",\n    r\"riss(bildung|e|es)?\\b\", r\"einriss\", r\"\\bruptur\", r\"zerreiss\", r\"\\blasion\",\n    r\"\\byirtik\", r\"\\byirtig\", r\"\\bkopma\\b\", r\"butunluk kaybi\", r\"\\brupturu\\b\",\n    r\"\\bpuknuce\", r\"\\bruptur\", r\"\\bprekid\\b\", r\"\\bpukotin\",\n    r\"ρηξη\", r\"ρηξις\", r\"ρηγμα\",\n    r\"руптура\", r\"разкъсв\", r\"разрив\", r\"скъсв\",\n)\n\nDEGEN = _rx(\n    r\"degenerat\", r\"\\bmucoid\\b\", r\"\\bmyxoid\\b\", r\"\\bfray\", r\"\\bfissur\",\n    r\"dejeneratif\", r\"\\bmukoid\\b\", r\"degenerativn\", r\"εκφυλ\", r\"дегенерат\",\n    r\"\\bμυξοειδ\", r\"\\bμυξωδ\",\n    r\"\\bmuco ?ide\\b\", r\"aufgefasert\",\n)\n\nINJURY = _rx(\n    r\"\\binjur\", r\"\\bsprain\", r\"\\blesion\", r\"\\blasion\", r\"\\bedema\\b\", r\"\\boedema\\b\",\n    r\"\\bodem\\b\", r\"\\bedem\\b\", r\"\\bοιδημα\", r\"\\bодем\", r\"\\bедем\", r\"\\bstrain\\b\",\n    r\"\\bhigh signal\\b\", r\"\\bsignal alteration\\b\", r\"\\bhiperintens\", r\"\\bhyperintens\",\n    r\"aumento de senal\", r\"alteracion de senal\", r\"cambio de senal\",\n    r\"\\bsignalanhebung\", r\"\\bsignalalteration\", r\"verhoogd signaal\", r\"sinyal artis\",\n    r\"αυξημενο σημα\", r\"повишен сигнал\",\n    r\"\\bthicken\", r\"\\bzadebljanje\\b\", r\"\\bverdikking\\b\", r\"\\bdistenzij\",\n    r\"\\blaksite\\b\", r\"\\blaxity\\b\", r\"\\bpartial\\b\", r\"\\bparcijaln\", r\"\\bparcial\",\n    r\"\\bpartiel\", r\"\\bpartiell\",\n)\n"},{"cell_type":"code","id":"code-08","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"ANAT = {\n    \"ACL\": _rx(\n        r\"anterior cruciate\", r\"\\bacl\\b\",\n        r\"cruzado anterior\", r\"\\blca\\b\",\n        r\"croise anterieur\",\n        r\"voorste kruisband\", r\"\\bvkb\\b\",\n        r\"vorderes kreuzband\", r\"vorderen kreuzband\", r\"vordere kreuzband\",\n        r\"on capraz\", r\"\\bocb\\b\",\n        r\"prednji krizni\", r\"prednjeg krizn\",\n        r\"προσθι[οα][^ ]* χιαστ\", r\"προσθιου χιαστου\", r\"χιαστο[^ ]* συνδεσμ\",\n        # \"χιαστοι και πλαγιοι συνδεσμοι\" separates the adjective from its noun, so the\n        # adjective stem has to stand alone. Greek marks cruciate with it unambiguously.\n        r\"\\bχιαστ\\w*\",\n        r\"предна кръстна\", r\"предната кръстна\",\n        # Plural, unqualified: reports routinely clear both cruciates in one clause\n        # (\"Ligamentos cruzados y colaterales dentro de limites normales\"), so the\n        # plural form has to match without a side qualifier or the whole clause is lost.\n        r\"cruciate ligaments\", r\"ligamentos cruzados\", r\"ligaments croises\",\n        r\"kruisbanden\", r\"kreuzbander\", r\"capraz baglar\", r\"krizn[a-z]* ligament[a-z]*\",\n        r\"χιαστοι συνδεσμ\", r\"χιαστων συνδεσμ\", r\"кръстните връзки\", r\"кръстни връзки\",\n    ),\n    \"MCL\": _rx(\n        r\"medial collateral\", r\"\\bmcl\\b\", r\"tibial collateral\",\n        r\"colateral medial\", r\"colateral interno\", r\"\\blcm\\b\",\n        r\"collateral medial\", r\"collateral interne\",\n        r\"mediale collaterale\", r\"binnenband\", r\"\\b(mediale|laterale) banden\\b\",\n        r\"\\bcollaterale banden\\b\",\n        r\"innenband\", r\"mediales? kollateral\",\n        r\"\\bic yan bag\", r\"medial kollateral\", r\"\\biyb\\b\",\n        r\"medijalni kolateraln\", r\"medijalnog kolateraln\",\n        r\"εσω πλαγι\", r\"εσωτερικο πλαγι\", r\"\\bπλαγι\\w* συνδεσμ\", r\"\\bπλαγιοι\\b\",\n        r\"медиален колатерал\", r\"вътрешна странична\", r\"\\bколатерал\\w*\",\n        # Same plural pattern as the cruciates.\n        # \"Ligamentos cruzados y colaterales\" separates the noun from its adjective, so\n        # the adjective has to stand alone as a cue.\n        r\"\\bcolaterales\\b\", r\"\\bcollateraux\\b\", r\"\\bcollateralen\\b\", r\"\\bkolateralni\\b\",\n        r\"collateral ligaments\", r\"ligamentos colaterales\", r\"ligaments collateraux\",\n        r\"collaterale banden\", r\"kollateralbander\", r\"seitenbander\", r\"yan baglar\",\n        r\"kolateraln[a-z]* ligament[a-z]*\", r\"πλαγιοι συνδεσμ\", r\"πλαγιων συνδεσμ\",\n        r\"колатерални връзки\", r\"страничните връзки\",\n    ),\n    \"Medial Meniscus\": _rx(\n        r\"medial meniscus\", r\"\\bmm\\b(?= tear)\", r\"medial menisc\",\n        r\"menisco medial\", r\"menisco interno\",\n        r\"menisque medial\", r\"menisque interne\",\n        r\"mediale meniscus\", r\"binnenmeniscus\",\n        r\"innenmeniskus\", r\"medialen? meniskus\", r\"innenmeniskushinterhorn\",\n        r\"medyal menisk\", r\"\\bic menisk\",\n        r\"medijalni meniskus\", r\"medijalnog meniskusa\", r\"medijalnom meniskusu\",\n        r\"εσω μηνισκ\", r\"μηνισκ[^ ]* του εσω\", r\"εσω διαμερισμα[^.]{0,40}μηνισκ\",\n        r\"медиалния менискус\", r\"медиален менискус\", r\"вътрешния менискус\",\n    ),\n    \"Lateral Meniscus\": _rx(\n        r\"lateral meniscus\", r\"lateral menisc\",\n        r\"menisco lateral\", r\"menisco externo\",\n        r\"menisque lateral\", r\"menisque externe\",\n        r\"laterale meniscus\", r\"buitenmeniscus\",\n        r\"aussenmeniskus\", r\"lateralen? meniskus\",\n        r\"lateral menisk\", r\"\\bdis menisk\",\n        r\"lateralni meniskus\", r\"lateralnog meniskusa\", r\"lateralnom meniskusu\",\n        r\"εξω μηνισκ\", r\"μηνισκ[^ ]* του εξω\", r\"εξω διαμερισμα[^.]{0,40}μηνισκ\",\n        r\"латералния менискус\", r\"латерален менискус\", r\"външния менискус\",\n    ),\n}\n\n# Osteoarthritis is rarely written as \"osteoarthritis\". It is written as cartilage loss,\n# chondropathy grade, joint space narrowing, or osteophytes - scoped to a compartment.\nOA_EVIDENCE = _rx(\n    r\"osteoarthrit\", r\"\\barthros\", r\"\\bgonarthros\", r\"\\bosteoarthros\",\n    r\"chondropath\", r\"chondromalac\", r\"condropat\", r\"condromalac\",\n    r\"cartilage loss\", r\"cartilage thinning\", r\"chondral (loss|defect|ulcer|thinning)\",\n    r\"osteophyt\", r\"osteofit\", r\"osteofyt\", r\"osteofito\", r\"osteophyten\",\n    r\"joint space narrowing\", r\"pinzamiento articular\",\n    r\"kikirdak kayb\", r\"kikirdak incelme\", r\"kondropati\", r\"kondral\",\n    r\"kraakbeen(lijden|verlies)\", r\"gonartrose\", r\"artrose\",\n    r\"knorpel(verlust|schaden|defekt)\", r\"arthrose\", r\"gonarthrose\",\n    r\"hrskavic\", r\"hondromalac\", r\"artroz\", r\"osteoartrit\",\n    r\"χονδρ[^ ]*παθ\", r\"αρθριτ\", r\"αρθρωσ\", r\"οστεοφυτ\",\n    r\"αρθρικου χονδρου\", r\"εξαλειψη του αρθρικου χονδρου\",\n    r\"артроз\", r\"хондропат\", r\"остеофит\", r\"хрущял[^.]{0,30}(изтън|увред|дефект)\",\n    r\"ulcera[s]? condral\", r\"cartilago[^.]{0,25}(perdida|adelgaz)\",\n    r\"icrs grade\", r\"outerbridge\",\n)\n\nCOMPARTMENT = {\n    \"Medial OA\": _rx(\n        r\"medial (femorotibial|tibiofemoral|compartment)\",\n        r\"compartimento femorotibial medial\", r\"femorotibial interno\",\n        r\"mediaal femorotibiaal\", r\"mediale femorotibial\",\n        r\"medial femorotibial\", r\"medialen kompartiment\", r\"innere[sn]? kompartiment\",\n        r\"medyal femorotibial\", r\"ic kompartman\", r\"medyal kompartman\",\n        r\"medijaln[^ ]* (femorotibi|odjelj|kompartm)\",\n        r\"εσω διαμερισμα\", r\"εσω κνημιαι\", r\"εσω μηριαι\",\n        r\"медиалн[^ ]* (компартм|отдел|тибиал|феморотиб)\",\n        r\"medial (femoral|tibial) (condyle|plateau)\", r\"condilo femoral medial\",\n        r\"medialen? (femurkondyl|tibiaplateau)\", r\"mediale femorale condyl\",\n    ),\n    \"Lateral OA\": _rx(\n        r\"lateral (femorotibial|tibiofemoral|compartment)\",\n        r\"compartimento femorotibial lateral\", r\"femorotibial externo\",\n        r\"lateraal femorotibiaal\", r\"laterale femorotibial\",\n        r\"lateral femorotibial\", r\"lateralen kompartiment\", r\"aussere[sn]? kompartiment\",\n        r\"lateral femorotibial\", r\"dis kompartman\", r\"lateral kompartman\",\n        r\"lateraln[^ ]* (femorotibi|odjelj|kompartm)\",\n        r\"εξω διαμερισμα\", r\"εξω κνημιαι\", r\"εξω μηριαι\",\n        r\"латералн[^ ]* (компартм|отдел|тибиал|феморотиб)\",\n        r\"lateral (femoral|tibial) (condyle|plateau)\", r\"condilo femoral lateral\",\n        r\"lateralen? (femurkondyl|tibiaplateau)\", r\"laterale femorale condyl\",\n    ),\n    \"PF OA\": _rx(\n        r\"patellofemoral\", r\"femoropatellar\", r\"femoropatelar\", r\"patelofemoral\",\n        r\"retropatellar\", r\"retrorotulian\", r\"\\btrochlea\", r\"\\btroclea\", r\"\\btroklea\",\n        r\"\\bpatella\\b\", r\"\\bpatellar\\b\", r\"\\brotulian\", r\"\\brotula\\b\", r\"\\bpatele\\b\",\n        r\"\\bpatellae?\\b\", r\"patellofemoraal\", r\"femoropatellair\",\n        r\"επιγονατιδ\", r\"μηροεπιγονατιδ\", r\"τροχιλ\",\n        r\"пател\", r\"феморопател\", r\"тролх\",\n        r\"anterior compartment\", r\"compartimento anterior\", r\"prednj[^ ]* odjeljk\",\n    ),\n}\n\n# Self-declaring findings: the term itself is the finding.\nDIRECT = {\n    \"Effusion\": _rx(\n        r\"\\beffusion\", r\"joint fluid\", r\"intra ?articular fluid\", r\"\\bhydrops\\b\",\n        r\"derrame articular\", r\"\\bderrame\\b\", r\"liquido articular\",\n        r\"epanchement\",\n        r\"gewrichtsvocht\", r\"\\bvocht\\b\", r\"\\bhydrops\\b\", r\"gewrichtseffusie\",\n        r\"gelenkerguss\", r\"\\berguss\\b\", r\"gelenksergu\",\n        # \"diz eklemi ici sivi miktari ... artmis\" and \"eklem icerisinde yaygin sivi\n        # artisi\" both occur; the noun takes a possessive suffix, so `eklem ` alone\n        # misses. Match the stem plus any suffix.\n        r\"eklem\\w* ic\\w* sivi\", r\"efuzyon\", r\"eklem sivisi\",\n        r\"sivi (miktari|artisi|birikimi)\", r\"sivi artis\", r\"\\bsivi\\b[^.]{0,25}artmis\",\n        r\"\\bizljev\", r\"\\bizliv\", r\"zglobn[^ ]* tekucin\", r\"\\bhidrops\\b\",\n        r\"αρθρικ[^ ]* υγρ\", r\"υγρου ενδαρθρικα\", r\"ενδαρθρικ[^ ]* υγρ\", r\"ποσοτητα υγρου\",\n        r\"ενδαρθρικ\", r\"αρθρικη συλλογη\", r\"υγρο στην αρθρωση\", r\"υγρου στην αρθρωση\",\n        r\"ставен излив\", r\"излив\", r\"ставна течност\", r\"синовиална течност\",\n    ),\n    \"Synovitis\": _rx(\n        r\"synovit\", r\"sinovit\", r\"synovial (thickening|proliferation|hypertroph)\",\n        r\"synovitis\", r\"synoviale? (verdikking|proliferatie)\",\n        r\"synovialitis\", r\"synovialis(verdickung|proliferation)\",\n        r\"sinovijalitis\", r\"sinovitis\", r\"zadebljanje sinovij\",\n        r\"υμενιτιδα\", r\"συνοβιτιδα\", r\"υμενικ[^ ]* υπερτροφ\", r\"αρθρικου υμεν\",\n        r\"синовит\", r\"синовиал[^ ]* (задебел|пролифер)\",\n        r\"verdikkingen van (het )?synovium\", r\"pannus\",\n    ),\n    \"Baker's\": _rx(\n        r\"baker\", r\"popliteal cyst\", r\"quiste popliteo\", r\"quistes popliteos\",\n        r\"kyste poplite\", r\"popliteale? cyst\", r\"poplitealzyste\", r\"bakerzyste\",\n        r\"popliteal kist\", r\"\\bbakerova\\b\", r\"poplitealn[^ ]* cist\",\n        r\"κυστη baker\", r\"πολυχωρη συνοβιακη κυστη\", r\"κυστη του baker\",\n        r\"киста на бейкър\", r\"бейкърова киста\", r\"поплитеална киста\",\n        r\"gastrocnemio ?semimembranos\", r\"gastrocnemius semimembranosus burs\",\n    ),\n    \"Contusion\": _rx(\n        r\"\\bcontusion\", r\"bone bruise\", r\"bone marrow (o?edema|contusion)\",\n        r\"\\bkontuz\", r\"medular bone o?edema\", r\"marrow o?edema\",\n        r\"contusion osea\", r\"edema oseo\", r\"edema de medula osea\",\n        r\"oedeme osseux\", r\"contusion osseuse\",\n        r\"botcontusie\", r\"botoedeem\", r\"beenmergoedeem\", r\"botmergoedeem\",\n        r\"knochenmarkodem\", r\"knochenodem\", r\"kontusion\", r\"bone bruise\",\n        r\"kemik kontuzyonu\", r\"kemik iligi odemi\", r\"kemik odemi\",\n        r\"kostani edem\", r\"edem kosti\", r\"kontuzij\",\n        r\"οστεομυελικ[^ ]* οιδημα\", r\"οστικο οιδημα\", r\"μυελικο οιδημα\",\n        r\"костномозъчен едем\", r\"костен едем\", r\"контузионен\",\n    ),\n    \"Fracture\": _rx(\n        r\"\\bfractur\", r\"\\bfract\\b\",\n        r\"\\bfractura\", r\"\\bfracturas\\b\",\n        r\"\\bfractuur\", r\"\\bbreuk\\b\",\n        r\"\\bfraktur\", r\"\\bbruch\\b\",\n        r\"\\bkirik\\b\", r\"\\bkirigi\\b\", r\"\\bkirik\\b\",\n        r\"\\bfraktur\", r\"\\bprijelom\", r\"impresijsk[^ ]* fraktur\",\n        r\"καταγμα\", r\"καταγματ\",\n        r\"фрактур\", r\"счупван\", r\"фисур\",\n        r\"insufficiency fracture\", r\"stress fracture\", r\"avulsion fracture\",\n        r\"subchondral fracture\", r\"subkondral kiri\",\n    ),\n}\n\n# Terms that look like a finding but are not the finding being scored.\nDECOY = {\n    # `no fracture` is deliberately absent: a decoy skips the clause, so listing it here\n    # turned the commonest English denial into silence, and the study then pulled on the\n    # fracture head with the weight of a report that never mentioned fractures at all.\n    # `microfractur` is a surgical procedure and `fracture risk` a prediction; both stay.\n    \"Fracture\": _rx(r\"microfractur\", r\"\\bfracture (risk|prophyla)\"),\n    \"Baker's\": _rx(r\"meniscal cyst\", r\"quiste meniscal\", r\"ganglion\"),\n}\n\nPAIRED = {\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\"}\nOA_TARGETS = {\"Medial OA\", \"Lateral OA\", \"PF OA\"}\n"},{"cell_type":"code","id":"code-09","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"STEM_MENISCUS = _rx(r\"menisc\\w*\", r\"menisk\\w*\", r\"μηνισκ\\w*\", r\"мениск\\w*\")\nSTEM_CRUCIATE = _rx(r\"cruciate\", r\"cruzado\", r\"croise\", r\"kruisband\", r\"kreuzband\",\n                    r\"capraz bag\\w*\", r\"krizn\\w*\", r\"χιαστ\\w*\", r\"кръстн\\w*\",\n                    r\"\\bacl\\b\", r\"\\bpcl\\b\", r\"\\blca\\b\", r\"\\blcp\\b\", r\"\\bvkb\\b\",\n                    r\"\\bhkb\\b\", r\"\\bocb\\b\", r\"\\bacb\\b\")\nSTEM_COLLATERAL = _rx(r\"collateral\\w*\", r\"colateral\\w*\", r\"kollateral\\w*\",\n                      r\"collaterale\\w*\", r\"kolateraln\\w*\", r\"yan bag\\w*\",\n                      r\"πλαγι\\w*\", r\"колатерал\\w*\", r\"странич\\w*\",\n                      r\"innenband\\w*\", r\"aussenband\\w*\", r\"binnenband\\w*\",\n                      r\"\\bmcl\\b\", r\"\\blcl\\b\", r\"\\blcm\\b\", r\"\\biyb\\b\")\n\nSIDE_MEDIAL = _rx(r\"\\bmedial\\w*\", r\"\\bmedyal\\w*\", r\"\\bmedijaln\\w*\", r\"\\bmediaal\\w*\",\n                  r\"\\bmediale\\w*\", r\"\\bintern[oa]\\w*\", r\"\\binterne\\w*\", r\"\\binnen\\w*\",\n                  r\"\\bic\\b\", r\"\\bunutarnj\\w*\", r\"\\bεσω\\w*\", r\"\\bεσωτερικ\\w*\",\n                  r\"\\bмедиал\\w*\", r\"\\bвътреш\\w*\", r\"\\btibial collateral\\b\",\n                  r\"\\bbinnen\\w*\", r\"\\bmediaal\\b\")\nSIDE_LATERAL = _rx(r\"\\blateral\\w*\", r\"\\bextern[oa]\\w*\", r\"\\bexterne\\w*\", r\"\\bdis\\b\",\n                   r\"\\blateraln\\w*\", r\"\\baussen\\w*\", r\"\\bbuiten\\w*\", r\"\\bεξω\\w*\",\n                   r\"\\bεξωτερικ\\w*\", r\"\\bлатерал\\w*\", r\"\\bвъншн\\w*\",\n                   r\"\\bfibular collateral\\b\", r\"\\bvanjsk\\w*\")\nSIDE_ANTERIOR = _rx(r\"\\banterior\\w*\", r\"\\bant\\b\", r\"\\bon\\b\", r\"\\bprednj\\w*\",\n                    r\"\\bvorder\\w*\", r\"\\bvoorste\\b\", r\"\\bπροσθι\\w*\", r\"\\bпредн\\w*\",\n                    r\"\\banteriyor\\w*\", r\"\\bavant\\b\", r\"\\bant[eé]rieur\\w*\")\n\n# The contrary of SIDE_ANTERIOR, needed only to stop a side-blind cruciate cue firing on\n# the posterior ligament. It is never used to assert a target - there is no PCL target -\n# so it is deliberately narrow: `posterior horn` is one of the commonest phrases in a\n# knee report and must not be read as a cruciate qualifier, which is why the guard below\n# tests proximity to the cruciate stem rather than presence in the clause.\nSIDE_POSTERIOR = _rx(r\"\\bposterior\\w*\", r\"\\bpost[eé]rieur\\w*\", r\"\\bposteriore\\w*\",\n                     r\"\\bhinter\\w*\", r\"\\bachterste\\b\", r\"\\barka\\b\", r\"\\bstraznj\\w*\",\n                     r\"\\bzadnj\\w*\", r\"\\bοπισθι\\w*\", r\"\\bзадн\\w*\", r\"\\bpostero\\w*\")\n\n# Fracture is the target whose stem varies most across the corpus.\nSTEM_FRACTURE = _rx(r\"fractur\\w*\", r\"fraktur\\w*\", r\"fractuur\\w*\", r\"\\bfract\\b\",\n                    r\"kiri[kgğ]\\w*\", r\"prijelom\\w*\", r\"lom kosti\", r\"\\bbreuk\\w*\",\n                    r\"\\bbruch\\w*\", r\"καταγμα\\w*\", r\"καταγματ\\w*\", r\"фрактур\\w*\",\n                    # NOT a bare `fissur\\w*`: \"fisuras condrales\" and \"full thickness\n                    # fissures in the articular cartilage\" describe cartilage, not bone.\n                    # The stem has to be anchored to a bone word to mean a fracture.\n                    r\"счупван\\w*\", r\"fisur\\w* (osea|oseas|kost)\", r\"fissur\\w* kost\")\n\nSTEM_OA_COMPARTMENT = _rx(r\"compartment\\w*\", r\"compartimento\\w*\", r\"compartiment\\w*\",\n                          r\"kompartman\\w*\", r\"kompartiment\\w*\", r\"odjelj\\w*\",\n                          r\"διαμερισμα\\w*\", r\"компартм\\w*\", r\"\\bотдел\\w*\",\n                          r\"femorotibial\\w*\", r\"femorotibiaal\\w*\", r\"tibiofemoral\\w*\",\n                          r\"femoro tibial\\w*\", r\"κνημιαι\\w*\", r\"μηριαι\\w*\",\n                          r\"femoral condyl\\w*\", r\"tibial plateau\\w*\",\n                          r\"condilo femoral\", r\"platillo tibial\", r\"tibiaplateau\\w*\",\n                          r\"femurkondyl\\w*\", r\"femoralne? kondil\\w*\",\n                          r\"tibijaln\\w* plato\", r\"femoral kondil\\w*\",\n                          r\"tibia plato\", r\"tibyal plato\")\n\n\ndef _distance(clause: str, stem_rx: re.Pattern, qual_rx: re.Pattern, window: int = 55):\n    \"\"\"Characters from the nearest stem to the nearest qualifier, or None if none is near.\n\n    Character windows rather than token windows, because word order differs: English\n    puts the side before the noun, Greek and Bulgarian often after, and Turkish\n    attaches it as a separate preceding adjective.\n\n    A distance rather than a yes. Presence is enough to decide that a qualifier applies\n    to a structure, but not enough to decide which of two qualifiers applies: a knee\n    report says \"anterior horn\" and \"posterior horn\" constantly, so any window wide\n    enough to catch a real side word also catches an unrelated one, and two rules that\n    both answer \"yes\" cannot be told apart. Comparing how far away they are can.\n    \"\"\"\n    best = None\n    for m in stem_rx.finditer(clause):\n        lo = max(0, m.start() - window)\n        hi = min(len(clause), m.end() + window)\n        for q in qual_rx.finditer(clause[lo:hi]):\n            qs, qe = lo + q.start(), lo + q.end()\n            d = 0 if qs < m.end() and qe > m.start() else \\\n                min(abs(m.start() - qe), abs(qs - m.end()))\n            best = d if best is None else min(best, d)\n    return best\n\n\ndef _near(clause: str, stem_rx: re.Pattern, qual_rx: re.Pattern, window: int = 55):\n    \"\"\"True if a stem match has a qualifier within `window` characters either side.\"\"\"\n    return _distance(clause, stem_rx, qual_rx, window) is not None\n    return False\n\n\n# concept -> (stem, side) pairs used in addition to the phrase lexicons above\nSTEM_RULES = {\n    \"ACL\": (STEM_CRUCIATE, SIDE_ANTERIOR),\n    \"MCL\": (STEM_COLLATERAL, SIDE_MEDIAL),\n    \"Medial Meniscus\": (STEM_MENISCUS, SIDE_MEDIAL),\n    \"Lateral Meniscus\": (STEM_MENISCUS, SIDE_LATERAL),\n    \"Medial OA\": (STEM_OA_COMPARTMENT, SIDE_MEDIAL),\n    \"Lateral OA\": (STEM_OA_COMPARTMENT, SIDE_LATERAL),\n}\n"},{"cell_type":"code","id":"code-10","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"SEV_LOW = _rx(\n    r\"\\bsmall\\b\", r\"\\bminimal\\b\", r\"\\btrace\\b\", r\"\\bmild\\b\", r\"\\bslight\\b\",\n    r\"\\btiny\\b\", r\"\\bscant\\b\", r\"\\bmimimal\\b\", r\"\\bdiscrete\\b\", r\"\\bfocal\\b\",\n    r\"\\bleve\\b\", r\"\\bminim\", r\"\\bpeque\", r\"\\bligero\\b\", r\"\\bescaso\\b\", r\"\\bdiscreto\\b\",\n    r\"\\bhafif\\b\", r\"\\bminimal\\b\", r\"\\baz miktarda\\b\", r\"\\bsilik\\b\",\n    r\"\\bmanja\\b\", r\"\\bmanji\\b\", r\"\\bblago\\b\", r\"\\bdiskretn\", r\"\\bmalo\\b\",\n    r\"\\bgering\", r\"\\bdiskret\", r\"\\bkleine?r?\\b\", r\"\\bwenig\\b\", r\"\\bzarte?\\b\",\n    r\"\\bbeperkte?\\b\", r\"\\bgeringe\\b\", r\"\\bweinig\\b\", r\"\\blichte?\\b\",\n    r\"\\bηπι\", r\"\\bμικρ\", r\"\\bελαχιστ\",\n    r\"\\bминимал\", r\"\\bлек\", r\"\\bмалк\", r\"\\bнеголям\",\n)\n\nSEV_HIGH = _rx(\n    r\"\\blarge\\b\", r\"\\bmarked\\b\", r\"\\bmassive\\b\", r\"\\bsevere\\b\", r\"\\bextensive\\b\",\n    r\"\\bmoderate\\b\", r\"\\bgross\\b\", r\"\\bsignificant\\b\", r\"\\babundant\\b\", r\"\\btense\\b\",\n    r\"\\bmoderad\", r\"\\bimportante\\b\", r\"\\bsevera?\\b\", r\"\\bmarcad\", r\"\\bcuantios\",\n    r\"\\bbelirgin\\b\", r\"\\byaygin\\b\", r\"\\bileri\\b\", r\"\\bciddi\\b\", r\"\\bbol\\b\",\n    r\"\\bopsezan\\b\", r\"\\bveliki\\b\", r\"\\bizrazit\", r\"\\bznacajn\", r\"\\bumjeren\",\n    r\"\\bausgepragt\", r\"\\bdeutlich\", r\"\\bmassiv\", r\"\\bmassig\", r\"\\bgross\",\n    r\"\\buitgebreid\", r\"\\bgevorderd\", r\"\\bveel\\b\", r\"\\bmatige?\\b\",\n    r\"\\bμετρι\", r\"\\bμεγαλ\", r\"\\bεκτεταμεν\", r\"\\bευμεγεθ\", r\"\\bσοβαρ\",\n    r\"\\bголям\", r\"\\bизразен\", r\"\\bзначим\", r\"\\bумерен\", r\"\\bобилен\",\n)\n\n# OA is often asserted for the whole joint rather than per compartment\n# (\"tricompartmental osteoarthritis\", \"gonarthrose\", \"incipient OA of all three\n# compartments\"). Those statements are evidence for all three OA targets.\nGLOBAL_OA = _rx(\n    r\"tri ?compartment\", r\"all three compartment\", r\"global(ised)? (oa|osteoarthrit)\",\n    r\"\\bgonarthros\", r\"\\bgonartros\", r\"\\bgonarthrose\", r\"\\bgonartrose\",\n    r\"osteoarthritis of the knee\", r\"artrosis (de |)(la )?rodilla\", r\"knee osteoarthrit\",\n    r\"\\bdiz osteoartrit\", r\"\\bgonartroz\", r\"artroza koljena\",\n    r\"οστεοαρθριτιδα\", r\"αρθριτιδα του γονατος\",\n    r\"артроза на колянната\", r\"гонартроз\",\n    r\"degenerative joint disease\", r\"\\bdjd\\b\",\n)\n\n# A bare \"bone marrow oedema\" is not a contusion when it sits under a cartilage\n# defect: subchondral oedema beneath a worn compartment is reactive degenerative signal,\n# and reading it as a bruise turns every osteoarthritic knee into a trauma case.\nDEGENERATIVE_MARROW = _rx(\n    r\"subchondral\", r\"subcondral\", r\"subkondral\", r\"supkondraln\", r\"subchondraln\",\n    r\"υποχονδρι\", r\"субхондрал\", r\"subchondrale?\",\n    r\"\\bcyst\", r\"\\bquist\", r\"\\bzyste\\b\", r\"\\bcistic\", r\"reactive\", r\"reactivo\",\n)\n\nTRAUMA = _rx(\n    r\"\\bbruise\\b\", r\"\\bcontusion\", r\"\\bkontuz\", r\"\\bcontusion osea\\b\",\n    r\"\\btrauma\", r\"\\bimpaction\\b\", r\"\\bpivot shift\\b\", r\"\\bkissing\\b\",\n    r\"\\bacute\\b\", r\"\\bagudo\\b\", r\"\\bakut\", r\"\\bpivot kaymasi\\b\",\n    r\"\\bcontusion osseuse\\b\", r\"\\bbone bruise\\b\", r\"\\bbotcontusie\\b\",\n    r\"\\bконтузион\", r\"\\bμωλωπ\", r\"\\bkontuzij\",\n)\n"},{"cell_type":"code","id":"code-11","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"def _polarity(clause: str, anchor_end: int) -> str:\n    \"\"\"Classify one clause as positive, negative or uncertain for a matched term.\n\n    Scope is the whole clause. Clause segmentation already keeps statements short, and\n    a window in characters mis-scopes badly across languages with different word orders -\n    Turkish puts its negator at the end of the sentence, English at the front.\n    \"\"\"\n    if UNCERTAIN.search(clause):\n        return \"uncertain\"\n    if NEGATION.search(clause):\n        return \"negative\"\n    if NORMALITY.search(clause):\n        # \"meniscus normal\" negates; \"normal ... but tear\" does not.\n        if TEAR.search(clause) or re.search(r\"\\bgrade [34]\\b\", clause):\n            return \"positive\"\n        return \"negative\"\n    return \"positive\"\n\n\nclass _Matcher:\n    \"\"\"Phrase lexicon first, stem+side proximity as the fallback.\n\n    Exposes `.search` so it drops into the same slot as a compiled pattern.\n    \"\"\"\n\n    def __init__(self, phrase_rx, stem=None, side=None, window=55, contrary=None):\n        self.phrase_rx = phrase_rx\n        self.stem = stem\n        self.side = side\n        self.window = window\n        self.contrary = contrary\n\n    def search(self, clause):\n        m = self.phrase_rx.search(clause)\n        if m is not None and not self._wrong_side(clause):\n            return m\n        if self.stem is not None and _near(clause, self.stem, self.side, self.window):\n            return self.stem.search(clause)\n        return None\n\n    def _wrong_side(self, clause):\n        \"\"\"True when the clause names the other member of this structure's pair.\n\n        Some cues in the lexicon are side-blind by design: Greek separates the adjective\n        from its noun (\"cruciate and collateral ligaments\"), so the bare adjective stem\n        has to stand alone or the clause is lost. That stem then also matches the\n        posterior cruciate and the lateral collateral, neither of which is a target here,\n        and a positive outranks every negative in the scorer - so one PCL clause was\n        enough to override an explicit \"the ACL is normal\".\n\n        The test is proximity to the structure's own stem, not presence in the clause.\n        \"Posterior horn of the medial meniscus\" appears in a large share of knee reports\n        and says nothing about a cruciate; only a qualifier sitting beside the ligament\n        word is one. A clause naming both sides keeps the match, because it does mention\n        this target.\n        \"\"\"\n        if self.contrary is None or self.stem is None:\n            return False\n        other = _distance(clause, self.stem, self.contrary, self.window)\n        if other is None:\n            return False\n        own = _distance(clause, self.stem, self.side, self.window)\n        # Not \"is the other side mentioned\" but \"is it the nearer of the two\". A clause\n        # reading \"tear of the posterior horn of the medial meniscus; the cruciate\n        # ligaments are intact\" mentions posterior, and under a presence test that was\n        # enough to suppress the cruciate cue - which is the opposite of the intent,\n        # since the clause does describe the ligaments. Ties go to keeping the match:\n        # a clause naming both sides does mention this one.\n        return own is None or other < own\n\n\n# Which cue, if it sits beside the structure's stem, means the clause is about the other\n# member of the pair. Only the two structures with a side-blind cue need one.\nCONTRARY = {\"ACL\": SIDE_POSTERIOR, \"MCL\": SIDE_LATERAL}\n\nANAT_MATCH = {\n    tgt: _Matcher(ANAT[tgt], *STEM_RULES[tgt], contrary=CONTRARY.get(tgt))\n    for tgt in PAIRED\n}\nCOMPARTMENT_MATCH = {\n    \"Medial OA\": _Matcher(COMPARTMENT[\"Medial OA\"], *STEM_RULES[\"Medial OA\"]),\n    \"Lateral OA\": _Matcher(COMPARTMENT[\"Lateral OA\"], *STEM_RULES[\"Lateral OA\"]),\n    \"PF OA\": _Matcher(COMPARTMENT[\"PF OA\"]),\n}\nDIRECT_MATCH = {\n    tgt: _Matcher(_rx(rx.pattern, STEM_FRACTURE.pattern) if tgt == \"Fracture\" else rx)\n    for tgt, rx in DIRECT.items()\n}\n\n\ndef _severity(clause: str) -> float:\n    \"\"\"Weight one positive mention by how emphatic the sentence is.\n\n    Ordered, not calibrated. A \"moderate effusion\" must outrank a \"trace effusion\" and\n    both must outrank silence; the absolute numbers do not matter to AUC.\n    \"\"\"\n    high = SEV_HIGH.search(clause) is not None\n    low = SEV_LOW.search(clause) is not None\n    if high and not low:\n        return 1.0\n    if low and not high:\n        return 0.45\n    return 0.75                       # unqualified mention\n\n\ndef _score_clauses(cls, anat_rx, path_rx=None, decoy_rx=None, context_penalty=None,\n                   context_bonus=None):\n    \"\"\"Accumulate graded evidence over clauses for one target.\n\n    Returns (score, confidence, n_pos, n_neg). Positives are graded by severity and by\n    optional context regexes; negatives only matter when nothing positive was found,\n    because reports assert normality for every structure they check.\n    \"\"\"\n    n_pos = n_neg = n_unc = 0\n    best = 0.0\n    for c in cls:\n        m = anat_rx.search(c)\n        if not m:\n            continue\n        if decoy_rx is not None and decoy_rx.search(c):\n            continue\n        if path_rx is not None and not path_rx.search(c):\n            if NORMALITY.search(c) and not NEGATION.search(c):\n                n_neg += 1\n            continue\n        pol = _polarity(c, m.end())\n        if pol == \"positive\":\n            n_pos += 1\n            w = _severity(c)\n            if context_penalty is not None and context_penalty.search(c):\n                w *= 0.45\n            if context_bonus is not None and context_bonus.search(c):\n                w = min(1.0, w * 1.35)\n            best = max(best, w)\n        elif pol == \"negative\":\n            n_neg += 1\n        else:\n            n_unc += 1\n            best = max(best, 0.30)\n\n    if n_pos or n_unc:\n        # 0.52 .. 0.95, ordered by the strongest single mention, nudged by repetition.\n        score = min(0.95, 0.50 + 0.42 * best + 0.03 * min(n_pos, 3))\n        conf = min(1.0, 0.55 + 0.15 * n_pos)\n    elif n_neg:\n        score = max(0.04, 0.20 - 0.04 * n_neg)\n        conf = min(0.9, 0.45 + 0.12 * n_neg)\n    else:\n        score, conf = 0.28, 0.05          # silence sits above asserted-negative\n    return score, conf, n_pos, n_neg\n\n\ndef extract(report: str) -> dict:\n    \"\"\"Extract twelve (score, confidence) pairs from one report.\"\"\"\n    cls = clauses(report)\n    out = {}\n    path_paired = _rx(TEAR.pattern, DEGEN.pattern, INJURY.pattern)\n\n    for tgt in TARGETS:\n        if tgt in PAIRED:\n            s, c, npos, nneg = _score_clauses(cls, ANAT_MATCH[tgt], path_paired)\n        elif tgt in OA_TARGETS:\n            s, c, npos, nneg = _score_clauses(cls, COMPARTMENT_MATCH[tgt], OA_EVIDENCE)\n        elif tgt == \"Contusion\":\n            # Reactive subchondral oedema under a cartilage defect is osteoarthritis,\n            # not a bruise. Explicit trauma wording pushes the other way.\n            s, c, npos, nneg = _score_clauses(cls, DIRECT_MATCH[tgt], None, DECOY.get(tgt),\n                                              context_penalty=DEGENERATIVE_MARROW,\n                                              context_bonus=TRAUMA)\n        else:\n            s, c, npos, nneg = _score_clauses(cls, DIRECT_MATCH[tgt], None, DECOY.get(tgt))\n        out[tgt] = s\n        out[tgt + \"__conf\"] = c\n        out[tgt + \"__npos\"] = npos\n        out[tgt + \"__nneg\"] = nneg\n\n    # --- cross-target corrections ------------------------------------------ #\n    # A whole-joint osteoarthritis statement is evidence for every compartment that was\n    # not separately assessed. Without this, \"incipient OA of all three compartments\"\n    # scores zero on all three OA targets.\n    g_hits = [c for c in cls if GLOBAL_OA.search(c) and _polarity(c, 0) == \"positive\"]\n    if g_hits:\n        gscore = 0.50 + 0.42 * max(_severity(c) for c in g_hits)\n        for tgt in OA_TARGETS:\n            if out[tgt + \"__npos\"] == 0 and out[tgt + \"__nneg\"] == 0:\n                out[tgt] = max(out[tgt], gscore * 0.92)\n                out[tgt + \"__conf\"] = max(out[tgt + \"__conf\"], 0.4)\n\n    # Synovitis is frequently visible on the images and absent from the text, so silence\n    # is weak evidence of absence here in a way it is not for other findings. Effusion is\n    # its most reliable textual proxy - the two share a mechanism - so a silent synovitis\n    # inherits a fraction of the effusion evidence instead of falling to the floor.\n    if out[\"Synovitis__npos\"] == 0 and out[\"Synovitis__nneg\"] == 0:\n        out[\"Synovitis\"] = max(out[\"Synovitis\"], 0.28 + 0.45 * (out[\"Effusion\"] - 0.28))\n\n    return out\n"},{"cell_type":"markdown","id":"md-12","metadata":{},"source":"## 3. Reading the acquisition\n\n`train_series.csv` describes each series with an anatomical plane and two binary flags,\n`Fluid_Sensitive` and `Fat_Suppression`. The names denote two physically independent\nproperties — and that independence is what the delivered columns do not have: across the\ntraining series the two agree on every row, so as given they carry one axis rather than\ntwo. That is the first reason to recover both from the header.\n\n*Fluid sensitivity* is a property of the **contrast weighting**, set by repetition time\n$T_R$ and echo time $T_E$:\n\n$$\n\\text{weighting} \\;=\\;\n\\begin{cases}\nT_1 & T_R \\lesssim 800\\ \\text{ms}\\\\\nT_2 & T_R \\gtrsim 800\\ \\text{ms},\\ T_E \\gtrsim 60\\ \\text{ms}\\\\\n\\text{PD} & T_R \\gtrsim 800\\ \\text{ms},\\ T_E \\lesssim 60\\ \\text{ms}\n\\end{cases}\n$$\n\nFluid is bright on $T_2$, intermediate on proton density, dark on $T_1$. Gradient echo\nbreaks the rule — its $T_R$ is short by design — so it is settled by `ScanningSequence`\nfirst. *Fat suppression* is a **preparation** applied on top of any weighting, and is what\nmakes marrow oedema conspicuous. Both are read from `SeriesDescription`, `SequenceName` and\n`ScanOptions` where the protocol names them, and from $T_R$ and $T_E$ where it does not.\n\n### Which sequences to show the model\n\nA knee is read in three planes because the structures run in different directions:\ncruciate ligaments obliquely, best seen sagittally; collateral ligaments and the meniscal\nbody coronally; patellar cartilage and the retinacula axially. Crossing plane with the two\nacquisition axes gives the slots below, chosen so each of the twelve findings has at least\none sequence that shows it well.\n\n| slot | plane | weighting | fat sat | what it carries |\n|---|---|---|---|---|\n| `SAG_FLUID_FS` | sagittal | PD / T2 | yes | meniscal tears, marrow oedema, effusion |\n| `COR_FLUID_FS` | coronal | PD / T2 | yes | collateral ligaments, meniscal body, oedema |\n| `AX_FLUID_FS` | axial | PD / T2 | yes | patellofemoral joint, synovium, effusion |\n| `SAG_FLUID_NOFS` | sagittal | PD / T2 | no | meniscal morphology at high contrast-to-noise |\n| `COR_T1` | coronal | T1 | no | marrow architecture, cartilage and bone outline |\n| `SAG_T1` | sagittal | T1 | no | anatomy, chronic change |\n\nA study rarely has all six; a per-slot presence mask carries the absences into the head,\nwhich §6 uses.\n"},{"cell_type":"code","id":"code-13","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"from __future__ import annotations\n\nimport os\n\nfor _v in (\"OMP_NUM_THREADS\", \"OPENBLAS_NUM_THREADS\", \"MKL_NUM_THREADS\"):\n    os.environ.setdefault(_v, \"4\")\n\nimport gc\nimport hashlib\nimport json\nimport re\nimport time\nimport traceback\nfrom concurrent.futures import ThreadPoolExecutor\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\n# The label extractor is defined in the cells above when this runs as a notebook. As a\n# plain script it is imported from the package source, so the two paths share one\n# definition rather than keeping a copy each.\n\nT0 = time.time()\nSEED = 2026\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\n\nTARGETS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n           \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n           \"Contusion\", \"Fracture\"]\n\n\n# The centre crop has to be smaller than the smallest field of view in the corpus or it\n# silently does nothing. Measured over every training series, the acquired field of view\n# (Rows x PixelSpacing) has median 160 mm and runs from 70 to 320: a 160 mm crop is\n# larger than the image in 60% of series and is skipped for all of them, which leaves\n# their physical scale unnormalised. 130 mm is below the field of view of 99.6% of\n# series and still contains the joint.\nCROP_MM = 130.0\n\n# Cache resolution. Everything downstream may downsample from this, so it is set by the\n# most demanding configuration rather than by the default one.\nCACHE_IMG = 336\nGROUP = 3                  # slices per encoder input, stacked as the three channels\nN_GROUP_MAX = 1\nCACHE_FRACTION = 0.45      # share of free memory the pixel cache may take\nCACHE_BUDGET_MAX_GB = 24.0 # hard ceiling regardless of what the machine reports\nCACHE_BUDGET_GB = 12.0     # only the fallback, for a machine with no /proc/meminfo\nTEST_SHARE = 0.30          # floor on the test corpus relative to the training one, since\n                           # the visible test split is a stub and the scored one is not\nHDR_THREADS = 16\nPIX_THREADS = 12\nORDER_THREADS = 32         # slice-ordering is latency-bound on the mount, not CPU-bound\n# Ceiling for the ordering pass. It has to be a ceiling because the pass is hundreds of\n# thousands of small reads over a network mount, so its duration is a property of the\n# mount rather than of the work, and varies between runs that do the same reading. It must not be a tight one: giving up leaves\n# those series in file order, which is uncorrelated with anatomy, and that degradation is\n# silent. So the ceiling sits well above what the pass ordinarily needs: its purpose is\n# to stop the pass consuming the whole run on a slow mount, not to trim the ordinary\n# case, and a ceiling tight enough to bind on a normal day would trade a silent\n# degradation for a saving the run does not need.\nORDER_BUDGET_S = 5400\n\n# Resolution is the axis under test. A feature of width d mm survives resampling only if\n# the pixel pitch is at most d/2, and the pitch here is set by the crop above rather than\n# by the acquired field of view: CROP_MM / P. At 224 px that is 0.58 mm, above the 0.5 mm\n# a 1 mm tear needs; at 336 px it is 0.39 mm and clears it. Both configurations read the\n# same cache, so the comparison isolates the resize.\nRUNS = [\n    {\"name\": \"r224\", \"img\": 224},\n    {\"name\": \"r336\", \"img\": 336},\n]\n\nEPOCHS = 10\nBATCH_STUDIES = 8          # a study is a bag of up to N_SLOT slot images\nAUG_ROT_DEG = 8.0          # rigid jitter; see augment() for why neither flip is used\nAUG_SCALE = 0.08\nAUG_SHIFT = 0.05\nAUG_INTENSITY = 0.10\nLAT_MIN_OFFSET_MM = 20.0   # inside this the side is not readable from geometry; see\n                           # side_from_geometry()\nSLICE_BAND = (0.20, 0.80)  # fraction of the ordered stack read_slot samples across\n\n# --- What a slice IS, as opposed to how many of them there are --------------- #\n#\n# A member is a function of the pixels it was fitted on, and img/crop_mm/slices/band do\n# not determine those pixels by themselves. Four further decisions do, none of them\n# visible in any shape:\n#\n#   order          which slice is the next one along the stack\n#   lat            which knees are mirrored, and on what evidence\n#   slot_fallback  whether a T1 slot may be filled from a series that is not T1\n#   decode_fill    what stands in for a slice that would not decode\n#\n# `native` is the reading derived in the sections below. `legacy` is the reading an\n# imported member was fitted under. A member read under the wrong one loads with every\n# shape matching, runs, and writes a plausible submission computed from the wrong image -\n# so the choice travels with the member and is part of the key that decides which members\n# can share a decode. The legacy rules are reproduced rather than corrected: correcting\n# them would hand that member pixels its weights never saw.\nRULES_NATIVE = {\"order\": \"normal\", \"lat\": \"centre\",\n                \"slot_fallback\": False, \"decode_fill\": \"nearest\"}\nRULES_LEGACY = {\"order\": \"dominant_axis\", \"lat\": \"corner_x\",\n                \"slot_fallback\": True, \"decode_fill\": \"zero\"}\nRULES = dict(RULES_NATIVE)\nLEGACY_LAT_OFFSET_MM = 5.0   # the dead zone the legacy laterality rule was fitted with\n\nLR_HEAD = 1e-3\nLR_BACKBONE = 8e-6         # the encoder is adapted, not retrained\nUNFREEZE_LAST = 6          # trainable transformer blocks, from the output end\nWEIGHT_DECAY = 0.02\nEVAL_BATCH = 8\nTIME_BUDGET = 8.0 * 3600\n\n# Six slots: three planes crossed with the acquisition axes. The fat-suppressed\n# fluid-sensitive series exist for nearly every study; the T1 and the non-suppressed\n# fluid-sensitive series are scarcer, which is what the presence mask is for.\nSLOTS_RECOVERED = [\n    (\"SAG_FLUID_FS\", \"Sagittal\", True, True),\n    (\"COR_FLUID_FS\", \"Coronal\", True, True),\n    (\"AX_FLUID_FS\", \"Axial\", True, True),\n    (\"SAG_FLUID_NOFS\", \"Sagittal\", True, False),\n    (\"COR_T1\", \"Coronal\", False, False),\n    (\"SAG_T1\", \"Sagittal\", False, False),\n]\n\n# The alternative: plane x the single axis the delivered flags carry, ignoring the\n# recovered weighting. Kept as\n# a switch so the choice of slot definition can be varied while everything else is held\n# fixed. Under this scheme a `Struct` slot mixes T1 series with non-fat-suppressed PD/T2\n# series, which carry very different tissue contrast.\nSLOTS_PUBLIC = [\n    (\"SAG_FLUID\", \"Sagittal\", None, True),\n    (\"COR_FLUID\", \"Coronal\", None, True),\n    (\"AX_FLUID\", \"Axial\", None, True),\n    (\"SAG_STRUCT\", \"Sagittal\", None, False),\n    (\"COR_STRUCT\", \"Coronal\", None, False),\n    (\"AX_STRUCT\", \"Axial\", None, False),\n]\n\nSLOT_SCHEME = os.environ.get(\"SLOT_SCHEME\", \"recovered\")\nSLOTS = SLOTS_PUBLIC if SLOT_SCHEME == \"public\" else SLOTS_RECOVERED\nN_SLOT = len(SLOTS)\n\n# How many 384-wide parts the per-slot feature is built from. The encoder emits one\n# vector per token; a slot feature is a fixed summary of that grid, and the summary an\n# imported member was fitted with carries a third part.\nPOOL_PARTS = {\"cls_mean\": 2, \"cls_mean_focal\": 3}\n\n# Which slots an imported member's attention is tilted toward, per diagnosis. Indices are\n# into SLOTS. This is a fixed table rather than a learned parameter, so it is part of that\n# member's definition and has to be reproduced exactly for its weights to mean anything.\nSLOT_PRIOR_TABLE = {\n    \"ACL\": (0, 3, 5), \"MCL\": (1, 4),\n    \"Medial Meniscus\": (0, 1, 3, 4), \"Lateral Meniscus\": (0, 1, 3, 4),\n    \"Medial OA\": (1, 4, 5), \"Lateral OA\": (1, 4, 5),\n    \"PF OA\": (0, 2, 5), \"Effusion\": (0, 2), \"Synovitis\": (0, 2),\n    \"Baker's\": (0,), \"Contusion\": (0, 1, 2), \"Fracture\": (0, 1, 2, 4, 5),\n}\nSLOT_PRIOR_STRENGTH = 0.55\n\nFATSAT_OPTS = {\"FS\", \"FATSAT\", \"FAT_SAT\", \"FSAT\"}\n_SEP = re.compile(r\"[_\\-.]\")\n_FATSAT_RX = re.compile(r\"\\bfs\\b|fatsat|fat sat|\\bstir\\b|\\bspair\\b|\\bspir\\b|\\bwe\\b|\"\n                        r\"water excit|\\btirm\\b|\\bsting\\b|\\bfatsup\\b\")\n_T1_RX = re.compile(r\"\\bt1\\b|\\bt1w\\b\")\n_T2_RX = re.compile(r\"\\bt2\\b|\\bt2w\\b\")\n_PD_RX = re.compile(r\"\\bpd\\b|\\bpdw\\b|proton|\\bdp\\b|dens\")\n"},{"cell_type":"code","id":"code-14","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"def log(msg):\n    print(f\"[{time.time() - T0:7.1f}s] {msg}\", flush=True)\n\n\ndef find_root():\n    for c in [Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\"),\n              Path(\"/kaggle/input/rsna-knee-abnormality-detection\"),\n              Path(\"data\"), Path(\".\")]:\n        if (c / \"test.csv\").is_file() and (c / \"test_series\").is_dir():\n            return c\n    # last resort: two-level scan, because the mount is nested one deeper than usual\n    base = Path(\"/kaggle/input\")\n    if base.is_dir():\n        for depth1 in sorted(p for p in base.iterdir() if p.is_dir()):\n            for cand in [depth1] + sorted(p for p in depth1.iterdir() if p.is_dir()):\n                if (cand / \"test.csv\").is_file():\n                    return cand\n    raise FileNotFoundError(\n        f\"competition mount not found (cwd {Path.cwd()}); expected a directory holding \"\n        f\"test.csv and test_series/\")\n\n\ndef find_dinov2(variant=\"small\"):\n    \"\"\"Locate a mounted DINOv2 checkpoint directory by variant name.\"\"\"\n    base = Path(\"/kaggle/input\")\n    if not base.is_dir():\n        return None\n    hits = []\n    for root, dirs, files in os.walk(base):\n        dirs[:] = [d for d in dirs if d not in (\"train_series\", \"test_series\")]\n        if \"config.json\" in files and \"dinov2\" in root.lower():\n            hits.append(Path(root))\n    for h in hits:\n        if variant in str(h).lower():\n            return h\n    return hits[0] if hits else None\n\n\nLABEL_COLS = TARGETS + [t + \"__conf\" for t in TARGETS]\n\n\nclass LabelSourceError(RuntimeError):\n    \"\"\"Raised when the labels did not come from where this run intended.\n\n    Every other failure in this file is better survived than reported: a run that dies\n    after the cache is built has spent the expensive half and scores nothing, so the\n    guard around `main` swallows it and leaves the benchmark file behind. This one is\n    the exception. Training on the weaker labels does not look like a failure - it\n    completes, writes a plausible submission, and differs only in a log line - so it has\n    to stop the run rather than be absorbed by a guard designed for crashes.\n    \"\"\"\n\n\ndef find_label_table():\n    \"\"\"Locate a mounted table of pre-read report labels, if one is attached.\n\n    The lexicon turns a report into labels by matching morphology, and its failure\n    mode is silence: on a phrasing it does not carry it emits no opinion rather than a\n    wrong one. Silence is measurable without any ground truth - for each (report,\n    finding) pair, did anything match? - and that measurement says the misses are\n    concentrated in particular languages rather than spread evenly, on findings a knee\n    report almost always comments on.\n\n    Enumerating morphology for nine languages is the wrong instrument for that. Reading\n    the sentence is the right one, and a language model reads it. Against the annotated\n    studies the difference is large and one-sided, so when such a table is mounted it is\n    preferred; when it is not, the lexicon runs and the pipeline is unchanged. Both paths\n    produce the same columns, so nothing downstream knows which one supplied them.\n    \"\"\"\n    base = Path(\"/kaggle/input\")\n    cands = []\n    if base.is_dir():\n        for root, dirs, files in os.walk(base):\n            dirs[:] = [d for d in dirs if d not in (\"train_series\", \"test_series\")]\n            cands += [Path(root) / f for f in files if f.startswith(\"report_labels\")\n                      and f.endswith(\".csv\")]\n    cands += [p for p in (Path(\"data/derived/report_labels_v2.csv\"),) if p.is_file()]\n    for c in cands:\n        try:\n            head = pd.read_csv(c, nrows=1)\n        except Exception:\n            continue\n        if \"StudyInstanceUID\" in head.columns and all(t in head.columns for t in TARGETS):\n            return c\n    return None\n\n\ndef label_mount_attached():\n    \"\"\"True when an input directory was attached that is meant to carry a label table.\n\n    The fallback below is deliberate and has to stay silent for a run with no table\n    attached, because that is the ordinary case for anyone reading this notebook. It\n    must not stay silent for the other case: a table was attached and could not be used.\n    Those two are indistinguishable from the labels alone - both end with the lexicon -\n    so they are separated here by whether the mount exists at all.\n    \"\"\"\n    base = Path(\"/kaggle/input\")\n    if not base.is_dir():\n        return False\n    return any(\"label\" in p.name.lower() for p in base.iterdir() if p.is_dir())\n\n\ndef read_labels(train_df):\n    \"\"\"Labels for every training study, from a mounted table or from the lexicon.\n\n    Studies the mounted table does not cover fall back to the lexicon rather than being\n    dropped, so a partial table degrades coverage instead of losing rows.\n    \"\"\"\n    n = len(train_df)\n    lab = pd.DataFrame([extract(r) for r in train_df[\"Report\"].fillna(\"\")])\n    lab[\"StudyInstanceUID\"] = train_df[\"StudyInstanceUID\"].values\n    lab = lab.set_index(\"StudyInstanceUID\")\n\n    src = find_label_table()\n    if src is None:\n        if label_mount_attached():\n            raise LabelSourceError(\n                \"LABEL SOURCE: a label dataset is mounted but no usable table was found \"\n                \"in it. Falling back to the lexicon here would train on the weaker \"\n                \"labels and say so only in a log line, so the run stops instead.\")\n        log(f\"LABEL SOURCE: lexicon, {n} studies (no table mounted)\")\n        return lab\n\n    tab = pd.read_csv(src).set_index(\"StudyInstanceUID\")\n    missing = [c for c in LABEL_COLS if c not in tab.columns]\n    if missing:\n        raise LabelSourceError(\n            f\"LABEL SOURCE: {src} is missing {len(missing)} expected columns \"\n            f\"(first: {missing[0]!r}). Refusing to fall back silently.\")\n    hit = lab.index.intersection(tab.index)\n    if not len(hit):\n        raise LabelSourceError(\n            f\"LABEL SOURCE: {src} shares no StudyInstanceUID with train.csv.\")\n    log(f\"LABEL SOURCE: {src.name} covers {len(hit)} of {n} studies, \"\n        f\"lexicon for the remaining {n - len(hit)}\")\n    lab.loc[hit, LABEL_COLS] = tab.loc[hit, LABEL_COLS].values\n    return lab\n\n\nROOT = find_root()\nlog(f\"input root: {ROOT}\")\n\n\nIMG = CACHE_IMG            # kept as the name the pixel reader and cache use\n\n\ndef available_gb():\n    \"\"\"Memory this machine will actually lend, read rather than assumed.\n\n    A hardcoded ceiling is a guess about a machine the author is not sitting at, and a\n    guess that is too low costs coverage silently while a guess that is too high ends the\n    run. The machine will say, so it is asked.\n    \"\"\"\n    try:\n        with open(\"/proc/meminfo\") as fh:\n            info = {k.strip(): v for k, v in\n                    (l.split(\":\", 1) for l in fh if \":\" in l)}\n        return int(info[\"MemAvailable\"].split()[0]) / 1024 ** 2\n    except Exception:\n        return CACHE_BUDGET_GB / CACHE_FRACTION      # fall back to the old constant\n\n\ndef plan_cache(n_study, n_test=0):\n    \"\"\"Choose how many slices per slot the memory the machine has will allow.\n\n    The cache is n_study x n_slot x slices x IMG^2 bytes. Coverage is the cheap axis -\n    linear - and resolution the expensive one, so when the budget binds it is the slice\n    count that gives way rather than the pixel grid. Deciding once, from the training\n    corpus size, keeps train and test caches on the same group layout.\n\n    Only a fraction of what is free is taken. The rest is not slack: the encoder, its\n    activations, the pinned batches and the frames all come out of the same pool, and the\n    cache is the one allocation big enough that overshooting it kills the run outright.\n    \"\"\"\n    avail = available_gb()\n    budget = min(avail * CACHE_FRACTION, CACHE_BUDGET_MAX_GB)\n    # Both caches are held at once, and the test half is what the visible run cannot\n    # show: here it is a handful of studies, and at scoring it is the whole hidden set.\n    # Sizing against the training corpus alone therefore passes every run that can be\n    # watched and overruns the one that counts.\n    n_total = n_study + max(n_test, int(TEST_SHARE * n_study))\n    per_slice = n_total * N_SLOT * IMG * IMG\n    afford = int(budget * 1024 ** 3 // max(per_slice, 1))\n    groups = max(1, min(N_GROUP_MAX, afford // GROUP))\n    log(f\"memory: {avail:.1f} GB available, {budget:.1f} GB to the cache; \"\n        f\"sizing for {n_study} train + {n_total - n_study} test studies \"\n        f\"-> {groups} group(s) of {GROUP} = {groups * GROUP} slices per slot\"\n        + (f\" (wanted {N_GROUP_MAX})\" if groups < N_GROUP_MAX else \"\"))\n    return groups\n\n\nN_GROUP = plan_cache(len(pd.read_csv(ROOT / \"train.csv\")),\n                     len(pd.read_csv(ROOT / \"test.csv\")))\nCACHE_SLICES = GROUP * N_GROUP\nlog(f\"cache layout: {N_GROUP} groups x {GROUP} slices = {CACHE_SLICES} per slot\")\n"},{"cell_type":"code","id":"code-15","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"HDR_TAGS = [\"SeriesDescription\", \"SequenceName\", \"ScanOptions\", \"ScanningSequence\",\n            \"RepetitionTime\", \"EchoTime\", \"Laterality\", \"PixelSpacing\", \"Rows\",\n            \"Columns\", \"RescaleSlope\", \"RescaleIntercept\",\n            # Position and orientation are read from the same header probe() already\n            # opens, so they cost nothing, and they are what recovers the side when the\n            # Laterality tag is absent - which it is for half the studies here.\n            \"ImagePositionPatient\", \"ImageOrientationPatient\"]\n\n\ndef _hdr_vec(s, n):\n    \"\"\"Parse a DICOM multi-value string as stored by probe(): floats joined by `|`.\"\"\"\n    if not isinstance(s, str):\n        return None\n    try:\n        v = [float(x) for x in s.split(\"|\")]\n    except ValueError:\n        return None\n    return np.array(v) if len(v) >= n else None\n\n\ndef side_from_geometry(h):\n    \"\"\"Study -> 'L' / 'R' / None, from where the image sits in the patient.\n\n    `Laterality` (0020,0060) is Type 2C and may legitimately be absent; in this corpus it\n    is missing on exactly half the studies, and the vendors it is missing from are whole\n    vendors rather than scattered series. A study with no tag is not a left knee, but the\n    normalisation upstream treats it as one, so half the corpus was never normalised and\n    the five side-defined targets - the two menisci, the two tibiofemoral compartments\n    and the medial collateral ligament - saw that axis reversed on a large minority of it.\n\n    The patient coordinate system fixes this without the tag: +x is the patient's left, so\n    the centre of a right knee sits at negative x. The centre is used rather than\n    `ImagePositionPatient` itself because that is the corner of the image, which is offset\n    by half a field of view - enough to change the sign on a knee near the midline.\n\n    The median over a study's series is what is thresholded, not a single series: probe()\n    reads one arbitrary slice per series, which on a sagittal stack can sit anywhere\n    across the joint. Studies whose centre falls near the midline are left unresolved\n    rather than guessed - measured against the tagged half, the rule is right 97% of the\n    time overall and no better than chance inside 20 mm.\n    \"\"\"\n    cx = {}\n    for r in h.itertuples(index=False):\n        ipp = _hdr_vec(getattr(r, \"ImagePositionPatient\", None), 3)\n        iop = _hdr_vec(getattr(r, \"ImageOrientationPatient\", None), 6)\n        ps = _hdr_vec(getattr(r, \"PixelSpacing\", None), 2)\n        rows, cols = getattr(r, \"Rows\", None), getattr(r, \"Columns\", None)\n        if ipp is None or iop is None or ps is None or not rows or not cols:\n            continue\n        try:\n            c = ipp[:3] + iop[:3] * ps[1] * float(cols) / 2 + iop[3:6] * ps[0] * float(rows) / 2\n        except (TypeError, ValueError):\n            continue\n        cx.setdefault(r.StudyInstanceUID, []).append(float(c[0]))\n    out = {}\n    for st, xs in cx.items():\n        m = float(np.median(xs))\n        out[st] = None if abs(m) < LAT_MIN_OFFSET_MM else (\"R\" if m < 0 else \"L\")\n    return out\n\n\ndef side_from_corner_x(h):\n    \"\"\"The laterality an imported member was fitted under.\n\n    It thresholds the median raw `ImagePositionPatient` x over a study's series. That is\n    the x of the image *corner*, not of its centre, so it differs from the rule above by\n    up to half a field of view - which is enough to reverse the sign on a knee scanned\n    near the midline. The dead zone is 5 mm rather than 20 mm, so it also commits on\n    studies the rule above leaves unresolved.\n\n    Neither difference changes a shape. Each one decides whether a study is mirrored, and\n    a study mirrored one way at training and the other at inference presents the five\n    side-defined targets with their axis reversed.\n    \"\"\"\n    out = {}\n    for st, g in h.groupby(\"StudyInstanceUID\"):\n        xs = []\n        for r in g.itertuples(index=False):\n            ipp = _hdr_vec(getattr(r, \"ImagePositionPatient\", None), 3)\n            if ipp is not None and np.isfinite(ipp).all():\n                xs.append(float(ipp[0]))\n        if not xs:\n            out[st] = None\n            continue\n        x = float(np.median(xs))\n        # DICOM patient coordinates are LPS: +x is the patient's left.\n        out[st] = None if abs(x) < LEGACY_LAT_OFFSET_MM else (\"R\" if x < 0 else \"L\")\n    return out\n\n\ndef lat_of(h, tag=\"\"):\n    \"\"\"Study -> 'L' / 'R' / None: the tag where it exists, geometry where it does not.\n\n    The tag is present on exactly half the studies here and is sometimes an empty\n    string rather than absent, which is not the same as NaN. Treating the other half\n    as left-sided is what `normalise_laterality` did by omission, so the geometry\n    fallback is not a refinement - it is the difference between normalising half the\n    corpus and normalising all of it.\n    \"\"\"\n    geo = side_from_corner_x(h) if RULES[\"lat\"] == \"corner_x\" else side_from_geometry(h)\n    d, n_tag, n_geo, n_none, n_disagree = {}, 0, 0, 0, 0\n    for st, g in h.groupby(\"StudyInstanceUID\"):\n        v = [str(x).strip().upper() for x in g[\"Laterality\"].dropna()]\n        if RULES[\"lat\"] == \"corner_x\" and \"ImageLaterality\" in g.columns:\n            # The legacy rule reads the second tag too, so a study tagged only there is\n            # resolved from the tag rather than from geometry.\n            v += [str(x).strip().upper() for x in g[\"ImageLaterality\"].dropna()]\n        v = [x[0] for x in v if x and x[0] in (\"L\", \"R\")]\n        side = v[0] if v else None\n        if side is not None:\n            n_tag += 1\n            if geo.get(st) is not None and geo[st] != side:\n                n_disagree += 1\n        else:\n            side = geo.get(st)\n            n_geo += side is not None\n            n_none += side is None\n        d[st] = side\n    log(f\"{tag}laterality: {n_tag} from the tag, {n_geo} from geometry, \"\n        f\"{n_none} unresolved; tag and geometry disagree on {n_disagree} \"\n        f\"({n_disagree / max(n_tag, 1):.1%} of the tagged)\")\n    return d\n\n\n\ndef probe(item):\n    split, study, series, path = item\n    row = {\"split\": split, \"StudyInstanceUID\": study, \"SeriesInstanceUID\": series,\n           \"dir\": path}\n    try:\n        files = sorted(e.name for e in os.scandir(path) if e.name.endswith(\".dcm\"))\n        row[\"files\"] = files\n        row[\"n_slices\"] = len(files)\n        if not files:\n            return row\n        ds = pydicom.dcmread(os.path.join(path, files[len(files) // 2]),\n                             stop_before_pixels=True, force=True)\n        for t in HDR_TAGS:\n            v = getattr(ds, t, None)\n            if v is None:\n                row[t] = None\n            elif isinstance(v, (list, tuple)) or type(v).__name__ == \"MultiValue\":\n                row[t] = \"|\".join(str(x) for x in v)\n            else:\n                row[t] = str(v)\n    except Exception as exc:\n        row[\"err\"] = str(exc)[:120]\n    return row\n\n\ndef walk(split):\n    \"\"\"Every series directory of a split, with one header read per series.\n\n    An absent split returns an empty frame *with the columns annotate expects*. Returning\n    a bare DataFrame looks like the same thing and is not: the next call indexes\n    `SeriesDescription` and raises KeyError, so the branch that exists to survive a\n    missing split is what turns it into a crash.\n    \"\"\"\n    base = ROOT / split\n    items = []\n    if not base.is_dir():\n        return pd.DataFrame(columns=[\"split\", \"StudyInstanceUID\", \"SeriesInstanceUID\",\n                                     \"dir\", \"files\", \"n_slices\"] + HDR_TAGS)\n    for study in os.scandir(base):\n        if study.is_dir():\n            for series in os.scandir(study.path):\n                if series.is_dir():\n                    items.append((split, study.name, series.name, series.path))\n    with ThreadPoolExecutor(max_workers=HDR_THREADS) as pool:\n        rows = list(pool.map(probe, items))\n    return pd.DataFrame(rows)\n\n\ndef annotate(df):\n    \"\"\"Recover fat suppression and pulse-sequence weighting from the header.\"\"\"\n    desc = (df[\"SeriesDescription\"].fillna(\"\") + \" \" + df[\"SequenceName\"].fillna(\"\"))\n    desc = desc.str.lower().str.replace(_SEP, \" \", regex=True)\n\n    opts = df[\"ScanOptions\"].fillna(\"\").str.upper().str.split(\"|\")\n    # GE writes SAT_GEMS for spatial saturation, so ScanOptions must be matched as\n    # exact tokens; a substring test on \"SAT\" fires on non-fat-sat series.\n    opts_fs = opts.apply(lambda ts: any(t.strip() in FATSAT_OPTS for t in ts))\n    df[\"fatsat\"] = desc.str.contains(_FATSAT_RX) | opts_fs\n\n    tr = pd.to_numeric(df[\"RepetitionTime\"], errors=\"coerce\")\n    te = pd.to_numeric(df[\"EchoTime\"], errors=\"coerce\")\n    gre = df[\"ScanningSequence\"].fillna(\"\").str.upper().str.contains(\"GR\")\n    t1, t2, pdw = desc.str.contains(_T1_RX), desc.str.contains(_T2_RX), desc.str.contains(_PD_RX)\n\n    df[\"weight\"] = np.where(t1 & ~t2 & ~pdw, \"T1\",\n                     np.where(t2 & ~pdw, \"T2\",\n                       np.where(pdw, \"PD\",\n                         np.where(gre, \"GRE\",\n                           np.where(tr < 800, \"T1\",\n                             np.where(te > 60, \"T2\",\n                               np.where(tr >= 800, \"PD\", \"UNK\")))))))\n    df[\"fluid\"] = np.isin(df[\"weight\"], [\"PD\", \"T2\"])\n    df[\"px\"] = pd.to_numeric(\n        df[\"PixelSpacing\"].fillna(\"\").str.split(\"|\").str[0].replace(\"\", np.nan),\n        errors=\"coerce\")\n    return df\n"},{"cell_type":"code","id":"code-16","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"def pick_slots(series_df, plane_map):\n    \"\"\"One series per slot per study.\n\n    Ties are broken toward the stack with the most slices: a thicker stack samples the\n    joint more densely, and the three-slice sampler below benefits from the margin.\n    \"\"\"\n    series_df = series_df.copy()\n    series_df[\"plane\"] = series_df[\"SeriesInstanceUID\"].map(plane_map)\n    out = {}\n    for study, g in series_df.groupby(\"StudyInstanceUID\"):\n        chosen = {}\n        for name, plane, fluid, fs in SLOTS:\n            sel = (g[\"plane\"] == plane) & (g[\"fatsat\"] == fs)\n            # fluid=None means \"do not condition on weighting\" - the public scheme,\n            # where the single provided flag stands in for both axes at once.\n            if fluid is not None:\n                sel &= (g[\"fluid\"] == fluid)\n            cand = g[sel]\n            # A slot with no series matching its predicate stays empty, and no substitute\n            # is admitted from a neighbouring predicate. Relaxing the weighting to fill a\n            # T1 slot would draw from the pool `SAG_FLUID_NOFS` selects from, since that\n            # pool is what remains once the weighting is dropped: over the training corpus\n            # it would put one series in two slots for 2383 of 4407 studies and leave 56%\n            # of the T1 slot holding PD or T2. The presence mask would then assert a\n            # sequence that was never acquired, and the per-diagnosis softmax of §6 would\n            # divide its attention across two identical slots, giving one acquisition\n            # about twice the weight it carries in a study that holds both. The mask is\n            # there to say a slot is absent, which is what an absent slot is.\n            if len(cand) == 0 and RULES[\"slot_fallback\"] and fluid is False:\n                # The relaxation the paragraph above rejects, reproduced because an\n                # imported member was fitted with its T1 slots filled this way: over half\n                # of that member's training studies had a T1 slot holding a series that\n                # is not T1. Leaving those slots empty would present it with a presence\n                # mask it never saw.\n                cand = g[(g[\"plane\"] == plane) & (~g[\"fatsat\"])]\n            if len(cand):\n                chosen[name] = cand.sort_values(\"n_slices\", ascending=False).iloc[0]\n        out[study] = chosen\n    return out\n"},{"cell_type":"markdown","id":"md-17","metadata":{},"source":"## 3b. What order the slices are in\n\nA series is a directory of files, and the obvious way to walk it is to sort the file names.\nThat is wrong here, and wrong in a way that produces no error. The file name is the SOP\nInstance UID, assigned to be unique rather than ordered, so sorting by it yields a sequence\nuncorrelated with anatomy — measured on a series from this corpus, the rank correlation\nbetween file-name order and physical position is $\\rho \\approx 0.01$. Three things\ndownstream depend on that order: three adjacent slices as three channels become three\nunrelated cross-sections; the middle of the stack becomes a random subset; and reversing\nthe order to normalise laterality does nothing at all.\n\nThe true order is recoverable exactly from geometry every slice carries.\n`ImageOrientationPatient` gives the in-plane axes $\\hat{r}_x, \\hat{r}_y$ and\n`ImagePositionPatient` the position $p$ of the first voxel, so\n\n$$\\hat{n} \\;=\\; \\hat{r}_x \\times \\hat{r}_y, \\qquad k \\;=\\; p \\cdot \\hat{n},$$\n\nand $k$ increases monotonically along the stack. Because $k$ is signed and expressed in\npatient coordinates, the stack has a fixed direction along the body's left-right axis —\nwhich is what the laterality normalisation of §5 reverses. `InstanceNumber` is the fallback\nwhere the geometry tags are missing; it usually tracks $k$ up to sign, but interleaved and\nmulti-echo acquisitions number slices in an order that is not the order they occupy in\nspace, and it is not signed in patient coordinates.\n"},{"cell_type":"markdown","id":"md-18","metadata":{},"source":"## 4. Sampling: how many millimetres one pixel is allowed to be\n\nA DICOM slice of $N \\times N$ pixels with spacing $s$ mm/pixel covers $Ns$ millimetres of\nanatomy. Both vary across this corpus, so a fixed-pixel resize hands the encoder images\nwhose physical scale differs by a factor of several. But scale normalisation is only half\nof it; the other half is a hard limit.\n\n**A feature narrower than two pixels does not survive the resize.** To represent a\nstructure of width $d$ millimetres the pixel pitch must satisfy $s_{\\text{eff}} \\le d/2$,\nthe Nyquist condition applied to the resampling grid. A meniscal tear is one to three\nmillimetres, so at $d = 1$ mm the pitch must be at most $0.5$ mm — and if it is not, no\ncapacity downstream recovers the signal, because it was destroyed before the first\nconvolution. This is a property of the resize, not of the network.\n\nCropping to a constant physical extent $L$ and resampling to $P$ pixels fixes the pitch:\n\n$$n \\;=\\; \\Big\\lfloor \\frac{L}{s} \\Big\\rceil \\ \\text{pixels}, \\qquad\ns_{\\text{eff}} \\;=\\; \\frac{L}{P}\\ \\ \\text{mm/pixel}, \\qquad\n\\text{token} \\;=\\; 14\\,s_{\\text{eff}}\\ \\ \\text{mm}.$$\n\nThe crop must be smaller than the smallest field of view or it silently does nothing: if\n$L/s$ exceeds the image width the crop cannot be taken and that series passes through\nunnormalised, with no error. $L = 130$ mm is below the acquired field of view of almost\nevery series here while still containing the joint. The resize target then follows from the\ntear width rather than convention — at $L = 130$ mm an input of $224$ gives $0.580$\nmm/pixel, above the bound for a 1 mm feature, while $336$ gives $0.387$ mm/pixel and puts a\n$14$-pixel patch token at $5.4$ mm.\n\n**Intensity needs the same treatment**, MR having no absolute scale. Each series is\nnormalised to its own 1st and 99th percentile — over the sampled stack rather than per\nslice, so slices keep their relative contrast, and percentiles rather than extremes, so one\nbright vessel does not compress everything else.\n"},{"cell_type":"code","id":"code-19","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"ORDER_TAGS = [(0x0020, 0x0032), (0x0020, 0x0037), (0x0020, 0x0013)]\n\n# Series in which at least one sampled slice would not decode. A list rather than a\n# counter because appending is atomic under the reader threads, and reported rather than\n# swallowed: unreported, a decode failure is indistinguishable from a black knee.\nDECODE_FAILED = []\n\n\ndef cache_tag(rules=None):\n    \"\"\"The name a decoded cache is stored under.\n\n    It has to name everything that decides the pixels, not only their dimensions. Two\n    configurations that agree on resolution, slice count, crop and band but disagree on\n    how a slice is chosen produce different arrays of identical shape - so a tag built\n    from the dimensions alone lets the second attach to the first one's file and train\n    against pixels it never asked for, with nothing anywhere reporting a mismatch.\n\n    A native reading keeps the plain name, so caches decoded before the rules existed\n    stay valid; anything else earns a suffix.\n    \"\"\"\n    r = dict(RULES if rules is None else rules)\n    t = (f\"{CACHE_IMG}px_{CACHE_SLICES}sl_{int(CROP_MM)}mm_\"\n         f\"{SLICE_BAND[0]:.2f}-{SLICE_BAND[1]:.2f}\")\n    if {k: r.get(k, v) for k, v in RULES_NATIVE.items()} != RULES_NATIVE:\n        t += \"_\" + hashlib.md5(json.dumps(r, sort_keys=True).encode()).hexdigest()[:6]\n    return t\n\n\ndef _natural_key(name):\n    return tuple(int(x) if x.isdigit() else x.lower()\n                 for x in re.split(r\"(\\d+)\", str(name)))\n\n\ndef _order_dominant_axis(rec):\n    \"\"\"The slice order an imported member was fitted under.\n\n    It sorts on the raw patient coordinate along whichever axis varies most across the\n    stack, rather than on the projection onto the slice normal. The two differ by a sign,\n    not by a formula: measured over this corpus every sagittal series has a slice normal\n    with n_x in [-1.00, -0.98], so p.n is the negative of the raw x this sorts on and the\n    two stacks come out exactly reversed. Because the band sampler truncates rather than\n    rounds, its nine indices are not symmetric about the middle, so nine slices drawn from\n    a twenty-six slice stack under one order share two with the other.\n\n    Missing geometry falls back to `InstanceNumber` and then to a natural sort of the file\n    name, both at the same 80% threshold the imported pipeline used.\n    \"\"\"\n    files, d = rec[\"files\"], rec[\"dir\"]\n    rows = []\n    for pos, f in enumerate(files):\n        ipp = inst = None\n        try:\n            ds = pydicom.dcmread(os.path.join(d, f), force=True, stop_before_pixels=True,\n                                 specific_tags=[\"ImagePositionPatient\", \"InstanceNumber\"])\n            raw = getattr(ds, \"ImagePositionPatient\", None)\n            if raw is not None and len(raw) >= 3:\n                c = np.asarray(raw[:3], dtype=np.float64)\n                if np.isfinite(c).all():\n                    ipp = c\n            n = getattr(ds, \"InstanceNumber\", None)\n            if n is not None:\n                inst = float(n)\n        except Exception:\n            pass\n        rows.append((f, ipp, inst, pos))\n\n    placed = [r for r in rows if r[1] is not None]\n    need = max(2, int(0.8 * len(rows)))\n    if len(placed) >= need:\n        xyz = np.stack([r[1] for r in placed])\n        axis = int(np.argmax(np.ptp(xyz, axis=0)))\n        spare = float(np.nanmedian(xyz[:, axis]))\n        rows.sort(key=lambda r: (float(r[1][axis]) if r[1] is not None else spare,\n                                 r[2] if r[2] is not None else float(\"inf\"), r[3]))\n    elif sum(r[2] is not None for r in rows) >= need:\n        rows.sort(key=lambda r: (r[2] if r[2] is not None else float(\"inf\"), r[3]))\n    else:\n        rows.sort(key=lambda r: _natural_key(r[0]))\n    return [r[0] for r in rows], True\n\n\ndef order_slices(rec):\n    \"\"\"Return the series' files sorted along the through-plane axis.\n\n    A DICOM file name here is a SOP Instance UID, which is assigned arbitrarily. Sorting\n    by it therefore produces an order uncorrelated with anatomy - measured over one\n    series, Spearman between file-name rank and physical position is 0.009, i.e. none.\n    Anything that assumes the file order means something is then operating on noise: the\n    three channels of a \"2.5D\" input are three unrelated views rather than neighbouring\n    slices, \"the middle of the stack\" is a random subset, and reversing slice order to\n    normalise laterality reverses nothing meaningful.\n\n    The physical order is recoverable exactly. Each slice carries its position in patient\n    coordinates and the in-plane axes; projecting the position onto the slice normal\n    gives a signed through-plane coordinate, monotonic along the stack:\n\n        n = r_x  x  r_y ,      k = p . n\n\n    `InstanceNumber` is the fallback. It usually tracks the projection up to sign, but\n    interleaved and multi-echo acquisitions need not number slices in the order they\n    occupy in space - but the projection is signed in patient\n    coordinates, which is what laterality normalisation needs.\n    \"\"\"\n    if RULES[\"order\"] == \"dominant_axis\":\n        return _order_dominant_axis(rec)\n    files, d = rec[\"files\"], rec[\"dir\"]\n    keyed = []\n    for f in files:\n        k = None\n        try:\n            ds = pydicom.dcmread(os.path.join(d, f), force=True, stop_before_pixels=True,\n                                 specific_tags=ORDER_TAGS)\n            iop = np.asarray(ds.ImageOrientationPatient, dtype=float)\n            ipp = np.asarray(ds.ImagePositionPatient, dtype=float)\n            k = float(np.dot(ipp, np.cross(iop[:3], iop[3:])))\n        except Exception:\n            try:\n                k = float(ds.InstanceNumber)\n            except Exception:\n                k = None\n        keyed.append((k, f))\n    if any(k is None for k, _ in keyed):\n        # A series with no usable geometry keeps its arbitrary order; that is worse than\n        # sorting but better than dropping the series, and it is logged as a count.\n        return files, False\n    return [f for _, f in sorted(keyed, key=lambda t: t[0])], True\n\n\ndef read_slot(rec, n_slice=None, out_size=None):\n    \"\"\"`n_slice` physically spread slices from one series, at `out_size` pixels.\n\n    Returns uint8 [n_slice, out, out] normalised per-series to its 1st-99th\n    percentile. Percentiles rather than min/max because MR intensity has no absolute\n    scale and a single bright vessel would otherwise compress the whole dynamic range.\n\n    Reading is the expensive half of this pipeline, so the caller reads once at the\n    largest configuration it needs and derives the smaller ones from the returned buffer\n    rather than re-reading.\n    \"\"\"\n    n_slice = GROUP if n_slice is None else n_slice\n    out_size = IMG if out_size is None else out_size\n    files, d, px = rec.get(\"ordered\") or rec[\"files\"], rec[\"dir\"], rec[\"px\"]\n    n = len(files)\n    if n == 0:\n        return None\n    # Spread the samples over a central band of the stack: the outermost slices of a knee\n    # series are mostly soft tissue outside the joint. The band is a constant rather than\n    # a literal because how much of the stack is worth reading depends on how many slices\n    # are being taken - at three the middle is all that fits, while at sixteen the ends\n    # are worth having, and a Baker cyst sits at the posteromedial end of a sagittal one.\n    lo, hi = int(SLICE_BAND[0] * (n - 1)), int(SLICE_BAND[1] * (n - 1))\n    idx = np.unique(np.linspace(lo, hi, n_slice).astype(int)) if hi > lo else np.array([n // 2])\n    while len(idx) < n_slice:\n        idx = np.append(idx, idx[-1])\n\n    planes = []\n    for i in idx[:n_slice]:\n        try:\n            ds = pydicom.dcmread(os.path.join(d, files[int(i)]), force=True)\n            a = ds.pixel_array.astype(np.float32)\n            sl = float(getattr(ds, \"RescaleSlope\", 1) or 1)\n            ic = float(getattr(ds, \"RescaleIntercept\", 0) or 0)\n            a = a * sl + ic\n        except Exception:\n            a = None                      # no shape is known here; see below\n        planes.append(a)\n\n    # A slice that would not decode has no shape of its own, and inventing one is how a\n    # single unreadable file erases a whole series: a substitute allocated at the resize\n    # target while the decoded slices are still native makes the shape check below take\n    # the substitute as the authority and zero the good slices with it, leaving a black\n    # slot that the presence mask still reports as acquired.\n    #\n    # A failure is instead filled from the nearest slice that did decode - the same\n    # convention the sampler already uses when the band holds fewer distinct slices than\n    # were asked for - and a series where nothing decodes is reported absent, which the\n    # mask can express, rather than black, which it cannot.\n    got = [k for k, p in enumerate(planes) if p is not None]\n    if RULES[\"decode_fill\"] == \"zero\":\n        # What an imported member was fitted with: a failure becomes a zero plane at the\n        # resize target, which the shape check below then propagates to the whole slot.\n        # It is the behaviour the paragraph above describes and rejects, kept here only\n        # because that member's weights were learned against slots blacked out this way.\n        if not got:\n            DECODE_FAILED.append(rec.get(\"SeriesInstanceUID\", d))\n        planes = [np.zeros((out_size, out_size), np.float32) if p is None else p\n                  for p in planes]\n        got = list(range(len(planes)))\n    if not got:\n        DECODE_FAILED.append(rec.get(\"SeriesInstanceUID\", d))\n        return None\n    if len(got) < len(planes):\n        DECODE_FAILED.append(rec.get(\"SeriesInstanceUID\", d))\n        for k, p in enumerate(planes):\n            if p is None:\n                planes[k] = planes[min(got, key=lambda j: abs(j - k))]\n\n    # Slices of one series can still differ in matrix size - multi-echo and some\n    # reformats do - and those are genuinely not stackable.\n    shp = planes[0].shape\n    planes = [p if p.shape == shp else np.zeros(shp, np.float32) for p in planes]\n    vol = np.stack(planes)\n\n    # constant physical extent, then resize: PixelSpacing varies 3.4x across the corpus\n    if px and np.isfinite(px) and px > 0:\n        want = int(round(CROP_MM / px))\n        h, w = shp\n        if 16 < want < min(h, w):\n            cy, cx = h // 2, w // 2\n            half = want // 2\n            vol = vol[:, max(0, cy - half):cy + half, max(0, cx - half):cx + half]\n\n    lo_v, hi_v = np.percentile(vol, [1, 99])\n    vol = np.clip((vol - lo_v) / max(hi_v - lo_v, 1e-6), 0, 1)\n\n    t = torch.from_numpy(np.ascontiguousarray(vol)).unsqueeze(0)\n    t = F.interpolate(t, size=(out_size, out_size), mode=\"bilinear\", align_corners=False)\n    # uint8, not float32. These buffers queue up between the reader threads and the\n    # encoder, and at this size a float32 slot-series is several megabytes. Intensity is\n    # already normalised into [0, 1] here, so eight bits cost nothing that a bilinear\n    # resize has not already cost, and the queue is a quarter the size.\n    return (t.squeeze(0) * 255).round().clamp(0, 255).to(torch.uint8)\n"},{"cell_type":"markdown","id":"md-20","metadata":{},"source":"## 5. Normalising left and right\n\nFour of the twelve targets — the two menisci and the medial and lateral tibiofemoral\ncompartments — are medial/lateral pairs, and a fifth, the medial collateral ligament, is\nnamed for the side it lies on. Medial and lateral are defined relative to the body's\nmidline, so which side of the *image* they fall on depends on which knee was scanned.\nUnless that is normalised, those five labels learn from an axis the model cannot observe.\n\nThe correction differs by plane. Coronally and axially the medial-lateral direction lies in\nthe image plane, so flipping the last axis maps one knee onto the other. Sagittally it is\nthe *slice* axis: each slice is unchanged by mirroring, and what differs is the order in\nwhich the stack traverses the joint, so the slice order is reversed instead.\n\n`Laterality` is a Type 2C attribute: it may legitimately be absent, and here it is absent\non half the studies — by whole vendors rather than scattered series. Leaving those alone\nsilently declares them left-sided, so every right knee among them enters the model\nmirrored. The patient coordinate system supplies the missing tag: position and orientation\nare recorded per image and $+x$ points to the patient's left, so the sign of the image\ncentre's $x$ says which knee this is.\n\n$$c \\;=\\; \\mathbf{p} \\;+\\; \\mathbf{r}\\,\\Delta_c \\frac{N_c}{2} \\;+\\; \\mathbf{d}\\,\\Delta_r \\frac{N_r}{2},\n\\qquad \\text{side} = \\begin{cases} \\text{right} & c_x < 0\\\\ \\text{left} & c_x > 0\\end{cases}$$\n\nwith $\\mathbf{p}$ the image position, $\\mathbf{r}$ and $\\mathbf{d}$ the row and column\ndirection cosines, and $\\Delta$ the pixel spacing. The centre is used rather than\n$\\mathbf{p}$ itself, which is a corner half a field of view away — enough to change the sign\non a knee scanned near the midline. The median over a study's series is thresholded, and a\nstudy centred within a short distance of the midline is left unresolved rather than guessed,\nbecause inside that band the sign is no better than chance.\n"},{"cell_type":"code","id":"code-21","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"def normalise_laterality(img, plane, lat):\n    \"\"\"Map every knee onto a left-knee convention.\n\n    Coronal and axial views mirror under a horizontal flip. Sagittal stacks are not\n    mirror images of each other - the slice order runs medial-to-lateral in opposite\n    directions - so the channel order is reversed instead.\n    \"\"\"\n    if lat != \"R\":\n        return img\n    if plane in (\"Coronal\", \"Axial\"):\n        return torch.flip(img, dims=[-1])\n    return torch.flip(img, dims=[0])\n"},{"cell_type":"markdown","id":"md-22","metadata":{},"source":"## 5b. Reading once, training many times\n\nThe cost of this pipeline is dominated by reading, not arithmetic. A study holds several\nseries and each series tens of slices, so a study is on the order of a hundred and fifty\nfiles. That is affordable once; it is not affordable once per epoch, and fine-tuning needs\nthe same pixels every epoch. So the slot images are decoded a single time into memory and\nheld as `uint8`:\n\n$$\\text{bytes} \\;=\\; N_{\\text{study}} \\times N_{\\text{slot}} \\times S \\times P^{2}$$\n\nwith $S$ slices kept per slot at $P$ pixels. The exponent on $P$ is what makes this a real\nconstraint rather than a detail — the cache grows with the *square* of resolution and only\nlinearly with slices, so coverage is the cheap axis and resolution the expensive one.\n\nThe cached slices form one three-channel encoder input per slot, and the layout generalises\nto several such groups per slot: training draws one per step, which doubles as augmentation\nalong the stack, and inference averages over them. What fixes the number of groups is a\nbudget rather than a capacity — the cache is allowed a fraction of the memory reported free,\ndeliberately below the whole of it, since it is the one allocation large enough that\novershooting ends the run rather than slowing it, and the encoder, its activations and the\nbuffers in flight come from the same pool. The size compared against that budget is the sum\nof *both* caches, training and test, because both are resident at once.\n\nA valid submission file is written before any of this begins and overwritten once real\npredictions exist, so the run always leaves a scoreable file behind.\n"},{"cell_type":"code","id":"code-23","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"# Where the geometric slice order may be remembered between runs. Unset on the platform,\n# because each run gets a fresh machine and there is nothing to remember; set off it,\n# where the same corpus is cached again at every resolution and slice count and the order\n# is a function of neither. It is opt-in so that the scored run's behaviour is decided by\n# the code rather than by whether a file happens to be lying about.\nORDER_CACHE = os.environ.get(\"RSNA_ORDER_CACHE\") or None\n\n\ndef build_cache(slot_map, plane_map, lat_map, tag):\n    \"\"\"Decode every (study, slot) once into an in-memory uint8 array.\n\n    Fine-tuning revisits the same pixels every epoch. Reading them from the mount each\n    time would make the epoch count a function of I/O rather than of learning, so they\n    are decoded once and held as bytes: intensity has already been normalised into\n    [0, 1], and eight bits cost nothing a bilinear resize has not already cost.\n\n    CACHE_SLICES positions are kept per slot, which the training loop reads as N_GROUP\n    groups of GROUP consecutive channels.\n    \"\"\"\n    studies = sorted(slot_map)\n    sidx = {s: i for i, s in enumerate(studies)}\n    cache = np.zeros((len(studies), N_SLOT, CACHE_SLICES, IMG, IMG), np.uint8)\n    mask = np.zeros((len(studies), N_SLOT), np.float32)\n    log(f\"{tag}: cache {cache.shape} = {cache.nbytes / 1024 ** 3:.1f} GB\")\n\n    jobs = [(st, k, plane, slot_map[st][name])\n            for st in studies\n            for k, (name, plane, _, _) in enumerate(SLOTS)\n            if name in slot_map[st]]\n    n_job = len(jobs)\n\n    # Ordering first, and as its own pass. It reads one header per slice of every chosen\n    # series - far more file opens than the decode that follows - and on a network mount\n    # that is latency, not work, so it gets its own wider pool.\n    t_ord = time.time()\n    n_slice_total = sum(len(j[3][\"files\"]) for j in jobs)\n    log(f\"{tag}: ordering {len(jobs)} slot-series ({n_slice_total} slice headers)\")\n    ok = done = 0\n    CHUNK_O = 1024\n\n    # A remembered order, when one is offered. The projection depends on the DICOM\n    # geometry alone, so it is the same at every resolution and every slice count, and\n    # it costs one header read per slice - the largest single cost in this pass. An entry\n    # is validated by the number of files present, so a tree that has changed under it is\n    # recomputed rather than trusted: order is derived data, and a stale entry would be\n    # invisible in the way that matters most.\n    seen = {}\n    if ORDER_CACHE and Path(ORDER_CACHE).is_file():\n        try:\n            import json as _json\n            seen = _json.loads(Path(ORDER_CACHE).read_text())\n        except (OSError, ValueError):\n            seen = {}\n        hit = 0\n        for _, _, _, rec in jobs:\n            e = seen.get(rec[\"SeriesInstanceUID\"])\n            if e and len(e[\"files\"]) == len(rec[\"files\"]):\n                rec[\"ordered\"] = e[\"files\"]\n                ok += int(e[\"good\"])\n                hit += 1\n        jobs = [j for j in jobs if \"ordered\" not in j[3]]\n        log(f\"{tag}: {hit} slot-series ordered from {ORDER_CACHE}, {len(jobs)} to read\")\n\n    with ThreadPoolExecutor(max_workers=ORDER_THREADS) as pool:\n        for c0 in range(0, len(jobs), CHUNK_O):\n            block = jobs[c0:c0 + CHUNK_O]\n            for (_, _, _, rec), (files, good) in zip(\n                    block, pool.map(lambda j: order_slices(j[3]), block)):\n                rec[\"ordered\"] = files\n                ok += int(good)\n                done += 1\n                if ORDER_CACHE:\n                    seen[rec[\"SeriesInstanceUID\"]] = {\"files\": files, \"good\": bool(good)}\n            # The ceiling is whichever comes first: the pass's own budget, or the share\n            # of what is left of the run that it may take. The second is what makes the\n            # first safe to set generously - a mount slow enough to matter cannot spend\n            # the training time, because the budget shrinks as the run does.\n            budget = min(ORDER_BUDGET_S, max(60.0, (TIME_BUDGET - (time.time() - T0)) * 0.35))\n            if time.time() - t_ord > budget:\n                log(f\"{tag}: ordering budget spent at {done}/{len(jobs)}; \"\n                    f\"the rest keep file order\")\n                break\n    if ORDER_CACHE and done:\n        import json as _json\n        _t = Path(ORDER_CACHE).with_suffix(\".tmp\")\n        _t.write_text(_json.dumps(seen))\n        _t.replace(Path(ORDER_CACHE))\n    log(f\"{tag}: ordered {ok}/{n_job} by geometry \"\n        f\"({n_job - ok} kept arbitrary) in {time.time() - t_ord:.0f}s\")\n\n    jobs = [(st, k, plane, slot_map[st][name])\n            for st in studies\n            for k, (name, plane, _, _) in enumerate(SLOTS)\n            if name in slot_map[st]]\n    log(f\"{tag}: decoding {len(jobs)} slot-series\")\n    n_failed_before = len(DECODE_FAILED)\n\n    CHUNK = 512\n    done = 0\n    with ThreadPoolExecutor(max_workers=PIX_THREADS) as pool:\n        for c0 in range(0, len(jobs), CHUNK):\n            block = jobs[c0:c0 + CHUNK]\n            for (st, k, plane, _), img in zip(\n                    block, pool.map(lambda j: read_slot(j[3], CACHE_SLICES, IMG), block)):\n                done += 1\n                if img is None:\n                    continue\n                cache[sidx[st], k] = normalise_laterality(img, plane,\n                                                          lat_map.get(st)).numpy()\n                mask[sidx[st], k] = 1.0\n            if done % 4096 < CHUNK:\n                log(f\"  {tag} {done}/{len(jobs)}\")\n            if time.time() - T0 > TIME_BUDGET:\n                log(f\"  {tag}: time budget reached during decode\")\n                break\n    n_failed = len(DECODE_FAILED) - n_failed_before\n    log(f\"{tag}: {int(mask.sum())}/{len(jobs)} slots filled\"\n        + (f\"; {n_failed} series had a slice that would not decode\" if n_failed else \"\"))\n    gc.collect()\n    return studies, cache, mask\n"},{"cell_type":"markdown","id":"md-24","metadata":{},"source":"## 6. Aggregating slots into twelve decisions\n\nA study arrives as up to six slot embeddings $x_s \\in \\mathbb{R}^{d}$ with a presence mask\n$m_s \\in \\{0,1\\}$. Pooling them identically would discard the reason the protocol has three\nplanes: each finding is read on particular sequences, and a mean over slots dilutes the one\ncarrying the evidence with five that do not.\n\nProject each slot, add a learned slot identity, give every diagnosis $o$ its own query\n$q_o \\in \\mathbb{R}^{H}$, and let it attend over the slots with absent ones masked out of\nthe softmax:\n\n$$h_s \\;=\\; \\phi(x_s) + e_s, \\qquad\n\\alpha_{o,s} \\;=\\; \\frac{\\exp\\!\\big(\\langle h_s, q_o\\rangle / \\sqrt{H}\\big)\\, m_s}\n{\\sum_{s'} \\exp\\!\\big(\\langle h_{s'}, q_o\\rangle / \\sqrt{H}\\big)\\, m_{s'}},$$\n\n$$c_o \\;=\\; \\sum_s \\alpha_{o,s}\\, h_s, \\qquad\n\\ell_o \\;=\\; \\langle c_o, w_o \\rangle + b_o .$$\n\nThe masked softmax renormalises over whatever the study actually contains, so a missing\naxial series shifts a diagnosis's attention onto the sequences that are present instead of\nfeeding it a zero vector.\n\n**The head is deliberately this small.** The label is attached to the *study*, so nothing\nin the supervision says which part of a study carries the finding, and a richer aggregation\nwould have no signal to learn that from. Where the supervision is coarse, so is the head.\n"},{"cell_type":"code","id":"code-25","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"class SlotHead(nn.Module):\n    \"\"\"Per-diagnosis attention over the slot embeddings of one study.\n\n    Each finding is read on particular sequences - cruciates sagittally, collateral\n    ligaments and the meniscal body coronally, patellar cartilage axially - so pooling\n    the slots identically would dilute the one that carries the evidence with the rest.\n\n    The aggregation is deliberately this simple. With a study-level label there is no\n    signal telling the model which part of a study matters, so extra attention\n    parameters below the slot level would have nothing to learn from and would spend\n    their capacity fitting noise.\n    \"\"\"\n\n    def __init__(self, dim, n_slot, n_out, hidden=256, p=0.2, prior=False):\n        super().__init__()\n        self.proj = nn.Sequential(nn.LayerNorm(dim), nn.Linear(dim, hidden), nn.GELU())\n        self.slot_emb = nn.Parameter(torch.randn(n_slot, hidden) * 0.02)\n        self.query = nn.Parameter(torch.randn(n_out, hidden) * 0.02)\n        self.drop = nn.Dropout(p)\n        self.out = nn.Linear(hidden, n_out)\n        self.hidden = hidden\n        # An imported member carries a fixed per-(diagnosis, slot) tilt on the attention\n        # logits, set from the anatomy table below rather than learned. It is a buffer, so\n        # it travels in the state dict and must exist for that member to load; exp(0.55)\n        # gives a preferred slot about 1.73x the weight of an unpreferred one, which\n        # biases the softmax without ever excluding a slot.\n        p_ = torch.zeros(n_out, n_slot)\n        if prior and n_slot == len(SLOTS) and n_out == len(TARGETS):\n            for t, slots in SLOT_PRIOR_TABLE.items():\n                if t in TARGETS:\n                    p_[TARGETS.index(t), list(slots)] = SLOT_PRIOR_STRENGTH\n        self.prior = prior\n        if prior:\n            self.register_buffer(\"slot_prior\", p_)\n\n    def forward(self, x, mask):\n        h = self.proj(x) + self.slot_emb\n        att = torch.einsum(\"bsh,oh->bos\", h, self.query) / self.hidden ** 0.5\n        if self.prior:\n            att = att + self.slot_prior.unsqueeze(0)\n        att = att.masked_fill(mask.unsqueeze(1) < 0.5, -1e4).softmax(-1)\n        ctx = self.drop(torch.einsum(\"bos,bsh->boh\", att, h))\n        return (ctx * self.out.weight.unsqueeze(0)).sum(-1) + self.out.bias\n"},{"cell_type":"code","id":"code-26","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"class Model(nn.Module):\n    \"\"\"Encoder plus head, trained end to end.\n\n    A study arrives as a bag of slot images. The bag is flattened for the encoder and\n    folded back before the head, so the encoder never sees the study structure and the\n    head never sees pixels.\n    \"\"\"\n\n    def __init__(self, backbone, dim, pool=\"cls_mean\", prior=False):\n        super().__init__()\n        self.backbone = backbone\n        self.pool = pool\n        self.head = SlotHead(dim * POOL_PARTS[pool], N_SLOT, len(TARGETS), prior=prior)\n        self.register_buffer(\"mean\", torch.tensor([0.485, 0.456, 0.406]).view(1, 3, 1, 1))\n        self.register_buffer(\"std\", torch.tensor([0.229, 0.224, 0.225]).view(1, 3, 1, 1))\n\n    def forward(self, imgs, mask, img_size=None):\n        B, S = imgs.shape[:2]\n        x = imgs.reshape(B * S, *imgs.shape[2:]).float().div_(255.0)\n        if img_size is not None and img_size != x.shape[-1]:\n            # The cache is held at the highest resolution any configuration needs; the\n            # rest downsample from it, so every configuration sees the same pixels\n            # through a different sampling grid rather than a different crop.\n            x = F.interpolate(x, size=(img_size, img_size), mode=\"bilinear\",\n                              align_corners=False)\n        x = (x - self.mean) / self.std\n        out = self.backbone(pixel_values=x).last_hidden_state\n        patch = out[:, 1:]\n        parts = [out[:, 0], patch.mean(1)]\n        if self.pool == \"cls_mean_focal\":\n            # The upper tail of each channel over the patch grid, taken per channel\n            # rather than by selecting whole patches: a finding occupies a small part of\n            # the field, so a plain mean over 256 patches dilutes it by two orders of\n            # magnitude, and this keeps the top eighth of each channel's responses.\n            k = max(1, patch.shape[1] // 8)\n            parts.append(patch.topk(k, dim=1).values.mean(1))\n        feat = torch.cat(parts, dim=1).reshape(B, S, -1)\n        return self.head(feat, mask)\n"},{"cell_type":"markdown","id":"md-27","metadata":{},"source":"### Why the encoder is trained rather than frozen\n\nA frozen self-supervised encoder is bounded by something no work downstream can reach.\nResolution, encoder size, slice coverage and slot aggregation change how much the model\nlooks and how closely, but none changes the vocabulary it looks *with*, so every one of those\naxes runs into the same ceiling — and that ceiling should be expected to bind here, since the\nencoder learned its features from natural images, where nothing resembles the signal a torn\nmeniscus makes on a proton-density sequence.\n\nSo the encoder is adapted, with two restraints. **Only the last blocks move** — early blocks\nof a vision transformer are generic edge and texture filters, late blocks are where\nsemantics live, and there may not be enough supervision here to improve the early ones while\nthere is certainly enough to damage them. **The encoder learns far more slowly than the\nhead** — the head is random at initialisation and has everything to learn, the encoder starts\nfrom a good solution and needs only to be moved off it, so a single learning rate would\neither leave the head untrained or destroy the encoder in the first few hundred steps.\n\nTargets remain the report-derived labels of §2, weighted by the per-finding confidence that\ncomes with them, so a report that never mentions a finding pulls weakly on that output\nrather than asserting a negative there. The studies carrying per-condition annotations are\nweighted above every derived row, since they are the only labels read from the images.\n"},{"cell_type":"code","id":"code-28","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"def build_model(unfreeze_last, source=None, variant=\"small\", pool=\"cls_mean\",\n                prior=False):\n    \"\"\"Load the encoder and open the last `unfreeze_last` blocks for training.\n\n    The early blocks of a self-supervised transformer are generic edge and texture\n    filters; the late blocks carry semantics. Opening only the late ones is the cautious\n    choice - there may not be enough supervision here to improve the early ones and there\n    is certainly enough to damage them - but how far the line should sit is a question\n    the corpus has to answer rather than the intuition.\n\n    `source` names where the weights come from. Left unset it is the attached model\n    directory, which is the only thing available here. It is a parameter so that a run\n    off the platform builds the same object from the same code rather than from a second\n    definition that has to be kept in step by hand.\n    \"\"\"\n    from transformers import AutoModel\n    p = source if source is not None else find_dinov2(variant)\n    if p is None:\n        raise FileNotFoundError(\"DINOv2 weights not attached\")\n    bb = AutoModel.from_pretrained(str(p))\n    n_layer = len(bb.encoder.layer)\n    for prm in bb.parameters():\n        prm.requires_grad = False\n    for blk in bb.encoder.layer[max(0, n_layer - unfreeze_last):]:\n        for prm in blk.parameters():\n            prm.requires_grad = True\n    for prm in bb.layernorm.parameters():\n        prm.requires_grad = True\n    dim = bb.config.hidden_size\n    trainable = sum(p.numel() for p in bb.parameters() if p.requires_grad)\n    log(f\"backbone: {n_layer} blocks, last {unfreeze_last} trainable \"\n        f\"({trainable / 1e6:.1f}M params), feature dim {dim * POOL_PARTS[pool]}\")\n    return Model(bb, dim, pool=pool, prior=prior)\n"},{"cell_type":"markdown","id":"md-29","metadata":{},"source":"## 6b. Reading weights instead of learning them\n\nNothing requires the scored run to be the run that learned the weights. A notebook may\nattach a dataset, weights are a dataset, and the part that genuinely cannot be done in\nadvance is the part depending on studies nobody has seen — reading them, and predicting.\nLearning here instead caps the model at what one accelerator fits inside the time limit,\nand spends that time again on every submission for a result that does not change. So when\na package of trained members is attached this section reads it; when none is, the sections\nbelow train one.\n\n**Members, not a model.** Each member carries the preprocessing it was fitted on. Two\nmembers fitted at different resolutions cannot share a decode; two fitted alike can. So\nmembers are grouped by the pixels they need, each group is decoded once, and a member added\nlater joins the list unchanged.\n\nWhat groups them is more than the resolution and the crop. Four further decisions settle\nwhat a slice *is* — which one comes next along the stack, which knees are mirrored, whether\na slot may be filled from a sequence that does not match its predicate, and what stands in\nfor a slice that will not decode — and none of them changes a single shape. A member fitted\nunder one reading and decoded under another therefore loads cleanly, runs, and returns a\nsubmission computed from the wrong image, so the reading travels with the member.\n\n**A member must prove it is the model it was.** Loading a state dictionary succeeds\nwhenever the shapes line up, and shapes line up across every difference that matters — a\nchanged normalisation, resize, or slice band. None of those raise. So each member carries\nthe answer it gave to a seeded question, recomputed before use; a mismatch stops the run\nrather than being averaged into a submission.\n\n**How a member is read.** Training draws one group of consecutive slices per step. At\ninference, looking more costs no extra decoding — the cache already holds $S$ slices, so one\ngroup from the middle is $1$ forward pass, the disjoint groups are $S/G$, and every\nconsecutive run of $G$ slices is $S-G+1$. Where the average is taken is free but not\nneutral: averaging logits then squashing is a geometric mean of odds, averaging\nprobabilities an arithmetic mean of risk, and they order studies differently. Members are\ncombined by rank, since §1 established order is all the metric reads.\n"},{"cell_type":"code","id":"code-30","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"FINGERPRINT_TOL = 2e-3\n\n\ndef fingerprint(model, dev, img_size, n_slot=None, group=None, seed=None):\n    \"\"\"The model's output on a fixed synthetic bag, as a portable identity.\n\n    Weights that are loaded but read through the wrong preprocessing produce predictions,\n    not errors. The submission is well formed, the log says nothing, and the difference is\n    a number no output of the run reveals. Scaling that never happens, or happens twice,\n    is enough on its own and changes no shape anywhere.\n\n    So a set of weights carries the answer it gave to a question with no data in it. The\n    input is generated from a seed rather than read, so it is the same on any machine, and\n    it is pushed through the whole forward path - the byte scaling, the ImageNet\n    normalisation, the resize, the encoder, the slot attention. Any of those differing\n    moves the output by order one. Numerics differing between two GPUs moves it by about\n    1e-5, which is why the tolerance sits between them rather than at zero.\n\n    This checks that the model computes what it computed when it was fitted. It cannot\n    check that the pixels reaching it are the right pixels; `read_slot` and the header\n    pass answer to their own tests.\n    \"\"\"\n    n_slot = N_SLOT if n_slot is None else n_slot\n    group = GROUP if group is None else group\n    seed = SEED if seed is None else seed\n    g = torch.Generator().manual_seed(seed)\n    imgs = torch.randint(0, 256, (2, n_slot, group, img_size, img_size),\n                         generator=g, dtype=torch.uint8).to(dev)\n    mask = torch.ones(2, n_slot, device=dev)\n    mask[1, -1] = 0.0                       # exercise the masked branch of the softmax\n    was_training = model.training\n    model.eval()\n    with torch.no_grad():\n        # float32 throughout: autocast would make the value depend on which device\n        # happened to run it, and the point of the number is that it does not.\n        out = model(imgs, mask, img_size).float().cpu().numpy()\n    if was_training:\n        model.train()\n    return out\n\n\ndef check_fingerprint(model, dev, img_size, expected, tol=FINGERPRINT_TOL, tag=\"\"):\n    \"\"\"Compare against a stored fingerprint; raise when the model is not the same map.\"\"\"\n    got = fingerprint(model, dev, img_size)\n    exp = np.asarray(expected, np.float32)\n    if got.shape != exp.shape:\n        raise WeightsError(f\"{tag}fingerprint shape {got.shape} != stored {exp.shape}: \"\n                           f\"the architecture is not the one these weights were fitted to\")\n    d = float(np.abs(got - exp).max())\n    if d > tol:\n        raise WeightsError(\n            f\"{tag}fingerprint differs by {d:.4g} (tolerance {tol:g}). The weights load \"\n            f\"but do not compute what they computed when fitted - preprocessing, \"\n            f\"resolution or architecture has moved between the two runs.\")\n    log(f\"{tag}fingerprint matches within {d:.2g}\")\n    return d\n\n\nclass WeightsError(RuntimeError):\n    \"\"\"Raised when attached weights cannot be trusted to be the ones that were fitted.\n\n    Deliberately fatal for the same reason as LabelSourceError: a run that predicts from\n    a mismatched model completes, writes a plausible submission, and differs from a\n    correct one only in a number no output of the run reveals.\n    \"\"\"\n\n\ndef find_weights(name=\"manifest.json\"):\n    \"\"\"Locate a mounted weights package, or return None if none is attached.\n\n    Same shape as `find_label_table`: the notebook must keep working for a reader who\n    attaches nothing, so absence is a path rather than an error. What must not be silent\n    is a package that is attached and unusable, and that is what `load_weights` refuses.\n    \"\"\"\n    import json\n    base = Path(\"/kaggle/input\")\n    if not base.is_dir():\n        return None\n    for root, dirs, files in os.walk(base):\n        dirs[:] = [d for d in dirs if d not in (\"train_series\", \"test_series\")]\n        if name not in files:\n            continue\n        # The manifest decides, not the filenames beside it. Testing for a naming\n        # convention makes the search agree with whatever the packager happened to call\n        # its files last, which is a second definition of what a package is.\n        try:\n            man = json.loads((Path(root) / name).read_text())\n        except (OSError, ValueError):\n            continue\n        if isinstance(man.get(\"members\"), list) and man[\"members\"]:\n            missing = [m[\"file\"] for m in man[\"members\"]\n                       if not (Path(root) / m[\"file\"]).is_file()]\n            if missing:\n                raise WeightsError(\n                    f\"{root} holds a manifest listing {len(man['members'])} members but \"\n                    f\"{len(missing)} of their files are absent (first {missing[0]!r})\")\n            return Path(root)\n    return None\n\n\n# How a member is read at inference. Overlapping windows over the slices the cache\n# already holds cost forward passes and no extra decoding, which is the cheap direction\n# to spend; and averaging probabilities rather than logits is an arithmetic mean of risk\n# rather than a geometric mean of odds, which orders studies differently. Both were\n# chosen by measuring them on the folds each member held out rather than by argument.\nTTA_OVERLAP = True\nTTA_POOL = \"prob\"\n\n\ndef window_starts(n_slice, group, overlap=None):\n    \"\"\"Where each TTA window begins.\"\"\"\n    overlap = TTA_OVERLAP if overlap is None else overlap\n    if overlap and n_slice >= group:\n        return list(range(n_slice - group + 1))\n    return [g * group for g in range(max(n_slice // group, 1))]\n\n\n@torch.no_grad()\ndef predict_member(model, cache, mask, idx, dev, img_size, group=None, pool=None,\n                   starts=None):\n    \"\"\"One member's predictions, averaged over its TTA windows.\n\n    `starts` is a parameter rather than always derived here so that a caller which has\n    measured what the remaining time affords can read a member over fewer windows. Passing\n    it is how the reduction becomes visible in one place instead of being re-derived.\n    \"\"\"\n    group = GROUP if group is None else group\n    pool = TTA_POOL if pool is None else pool\n    starts = window_starts(cache.shape[2], group) if starts is None else list(starts)\n    if not starts:\n        raise ValueError(\"predict_member was given no windows to average over\")\n    model.eval()\n    out = []\n    for b in range(0, len(idx), EVAL_BATCH):\n        sel = idx[b:b + EVAL_BATCH]\n        m = torch.from_numpy(mask[sel]).to(dev)\n        acc = None\n        for st in starts:\n            rows = torch.from_numpy(\n                np.ascontiguousarray(cache[sel, :, st:st + group])).to(dev)\n            with torch.autocast(\"cuda\", enabled=dev.type == \"cuda\"):\n                z = model(rows, m, img_size).float()\n            v = z if pool == \"logit\" else torch.sigmoid(z)\n            acc = v if acc is None else acc + v\n        v = acc / len(starts)\n        out.append((torch.sigmoid(v) if pool == \"logit\" else v).cpu().numpy())\n    return np.concatenate(out) if out else np.zeros((0, len(TARGETS)), np.float32)\n\n\ndef infer_from_package(path, dev):\n    \"\"\"Predict the test split from an attached package of trained members.\n\n    The alternative below - learning the weights inside the scored run - spends the whole\n    allowance on the training corpus every time the notebook is submitted, and caps the\n    model at what nine hours on one accelerator can fit. Neither is necessary: a notebook\n    may attach a dataset, and weights are a dataset. What the scored run then does is the\n    part that cannot be done in advance, because the studies are not known in advance.\n\n    Members are grouped by the pixels they need. Two members fitted at different\n    resolutions are different functions of the same study and cannot share a decode; two\n    fitted alike can, and that is the whole reason the grouping exists rather than a\n    decode per member.\n    \"\"\"\n    import json\n    man = json.loads((Path(path) / \"manifest.json\").read_text())\n    members = man[\"members\"]\n    members = members[-5:]  # V19-bot5\n    log(f\"V19-bot5: using {len(members)} members\")\n\n    test_df = pd.read_csv(ROOT / \"test.csv\")\n    test_series = pd.read_csv(ROOT / \"test_series.csv\")\n    plane_map = dict(zip(test_series[\"SeriesInstanceUID\"],\n                         test_series[\"Anatomical_Plane\"]))\n    hte = annotate(walk(\"test_series\"))\n    log(f\"test header pass: {len(hte)} series\")\n\n    groups = {}\n    for m in members:\n        groups.setdefault(m[\"pixel_group\"], []).append(m)\n\n    per_member = []\n    fixed_s = per_win_s = None\n    for gi, (key, gm) in enumerate(groups.items(), 1):\n        cfg = json.loads(key)\n        adopt_config_globals(cfg)\n        log(f\"decode group {gi}/{len(groups)}: {cfg['img']}px x {cfg['slices']} slices, \"\n            f\"crop {cfg['crop_mm']} mm -> {len(gm)} member(s)\")\n        st_te, Cte, Mte = build_cache(pick_slots(hte, plane_map), plane_map,\n                                      lat_of(hte, \"test \"), f\"test g{gi}\")\n        idx = np.arange(len(st_te))\n\n        # What the remaining time affords. A member costs a fixed part - loading the\n        # checkpoint, building the encoder, checking the fingerprint - plus a part that\n        # scales with the number of TTA windows it is read over. Only the second is worth\n        # trading, and only the first member can measure either, because the size of the\n        # test set is not something the notebook is told.\n        #\n        # Both parts are timed separately. Dividing a member's whole wall time by its\n        # window count prices the fixed part as if it scaled, so every reduction inflates\n        # the estimate that caused it and the estimate runs away from the truth.\n        #\n        # Windows are surrendered before members. A dropped window costs one view of a\n        # study that other windows also see; a dropped member costs an independent vote,\n        # which is what an ensemble is made of. The question asked before each member is\n        # whether ONE more fits - not whether all the rest do, which would abandon an\n        # ensemble that could still have run most of itself.\n        starts = window_starts(Cte.shape[2], GROUP)\n        order = sorted(gm, key=lambda m: -(m.get(\"holdout\") or 0))\n        left_after = sum(len(g) for j, (_, g) in enumerate(groups.items(), 1) if j > gi)\n        for k, m in enumerate(order):\n            left = TIME_BUDGET - (time.time() - T0)\n            remaining = (len(order) - k) + left_after\n            if fixed_s is not None and per_win_s is not None:\n                # Leave a tenth of what is left unspent: the estimate comes from one\n                # member on one machine, and a submission written late is not written.\n                afford = max(left * 0.9, 0.0)\n                need = fixed_s + len(starts) * per_win_s\n                if need * remaining > afford:\n                    room = afford / max(remaining, 1)\n                    n_win = int((room - fixed_s) / per_win_s) if per_win_s > 0 else 0\n                    n_win = max(1, min(len(starts), n_win))\n                    if fixed_s + per_win_s > afford:\n                        log(f\"  {left / 60:.0f} min left: stopping after {k} of \"\n                            f\"{len(order)} in this decode group; not one more member \"\n                            f\"fits, at any window count\")\n                        break\n                    if n_win < len(starts):\n                        mid = (len(starts) - n_win) // 2\n                        log(f\"  {left / 60:.0f} min left, {remaining} member(s) to go: \"\n                            f\"{n_win} window(s) each instead of {len(starts)}\")\n                        starts = starts[mid:mid + n_win]\n            t0 = time.time()\n            ck = torch.load(Path(path) / m[\"file\"], map_location=\"cpu\",\n                            weights_only=False)\n            model = build_model(int(m[\"config\"][\"unfreeze_last\"]),\n                                variant=m[\"config\"][\"variant\"],\n                                pool=m[\"config\"].get(\"pool\", \"cls_mean\"),\n                                prior=bool(m[\"config\"].get(\"prior\", False))).to(dev)\n            model.load_state_dict(ck[\"model\"])\n            check_fingerprint(model, dev, IMG, ck[\"fingerprint\"], tag=f\"{m['id']}: \")\n            t_ready = time.time()\n            p = predict_member(model, Cte, Mte, idx, dev, IMG, starts=starts)\n            per_member.append({\"id\": m[\"id\"], \"ids\": st_te, \"pred\": p,\n                               \"holdout\": m.get(\"holdout\")})\n            fixed_s = t_ready - t0\n            per_win_s = (time.time() - t_ready) / max(len(starts), 1)\n            log(f\"  {m['id']} fold {m['fold']}: predicted {len(idx)} studies over \"\n                f\"{len(starts)} window(s) in {time.time() - t0:.0f}s\")\n            del model, ck\n            gc.collect()\n            if dev.type == \"cuda\":\n                torch.cuda.empty_cache()\n        del Cte, Mte\n        gc.collect()\n\n    # Rank rather than probability, because the metric reads order and two members\n    # calibrated differently would otherwise have unequal say. Every member covers every\n    # test study, so the mean is over the same set each time and needs no weighting to be\n    # comparable; weighting by a fold's holdout would import into the test set a\n    # difference measured on a few hundred training studies.\n    all_ids = sorted({s for m in per_member for s in m[\"ids\"]})\n    pos = {s: i for i, s in enumerate(all_ids)}\n    acc = np.zeros((len(all_ids), len(TARGETS)), np.float64)\n    for m in per_member:\n        r = pd.DataFrame(m[\"pred\"]).rank(pct=True).to_numpy()\n        acc[[pos[s] for s in m[\"ids\"]]] += r\n    acc /= max(len(per_member), 1)\n\n    sub = write_submission(acc, all_ids, test_df, \"submission.csv\")\n    log(f\"submission.csv = rank mean of {len(per_member)} member(s); {sub.shape}; \"\n        f\"nulls {int(sub[TARGETS].isna().sum().sum())}\")\n    return sub\n\n\ndef adopt_config_globals(cfg):\n    \"\"\"Point the pixel path at what one group of members was fitted on.\"\"\"\n    global IMG, CACHE_IMG, GROUP, CACHE_SLICES, N_GROUP, CROP_MM, SLICE_BAND, RULES\n    CACHE_IMG = IMG = int(cfg[\"img\"])\n    GROUP = int(cfg[\"group\"])\n    CACHE_SLICES = int(cfg[\"slices\"])\n    N_GROUP = max(CACHE_SLICES // GROUP, 1)\n    CROP_MM = float(cfg[\"crop_mm\"])\n    SLICE_BAND = tuple(float(x) for x in cfg[\"band\"])\n    # The four decisions that change what a slice is. A member fitted under one reading\n    # and decoded under another gets pixels its weights never saw, with every shape\n    # still agreeing, so an unrecognised name is refused rather than defaulted.\n    rules = cfg.get(\"rules\") or RULES_NATIVE\n    unknown = {k: v for k, v in rules.items()\n               if k not in RULES_NATIVE\n               or v not in (RULES_NATIVE[k], RULES_LEGACY[k])}\n    if unknown:\n        raise WeightsError(f\"the members record pixel rules this pipeline cannot \"\n                           f\"reproduce: {unknown}\")\n    RULES = {**RULES_NATIVE, **rules}\n    if [s[0] for s in SLOTS] != list(cfg[\"slots\"]):\n        raise WeightsError(\n            f\"the members were fitted on slots {cfg['slots']} and this pipeline defines \"\n            f\"{[s[0] for s in SLOTS]}; a weight would be read against the wrong slot\")\n"},{"cell_type":"code","id":"code-31","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"def take_group(cache_rows, g):\n    \"\"\"Slice GROUP consecutive channels out of the cached slices.\"\"\"\n    return cache_rows[:, :, g * GROUP:(g + 1) * GROUP]\n\n\ndef augment(imgs):\n    \"\"\"A small rigid jitter and an intensity scale, applied to a whole bag at once.\n\n    Neither flip is available here, and for different reasons. A horizontal flip would\n    reintroduce the nuisance axis that the laterality normalisation removed - it would\n    undo, once per batch, what the header pass was run to establish.\n\n    A vertical flip is not a nuisance axis at all. A knee is acquired in a canonical\n    orientation, and no study in this corpus looks like its own vertical mirror. An\n    augmentation is meant to cover directions along which the label does not change; this\n    one moves the input off the distribution the encoder will be asked about, which is a\n    different thing. Where a finding sits in the frame is also information rather than\n    noise - a Baker cyst is identified by lying in the popliteal fossa, not by its\n    appearance alone.\n\n    What is left is jitter that no label depends on: a few degrees of rotation, a few\n    per cent of scale and translation. That still prevents memorising the exact framing,\n    which is what an augmentation is for, while leaving the anatomy where it was.\n    \"\"\"\n    # A bag arrives as [study, slot, GROUP, IMG, IMG]: five axes, not four. The warp is\n    # a 2-D operation, so the two leading axes are folded together and restored after -\n    # every slot image is an independent acquisition and gets its own jitter.\n    lead = imgs.shape[:-3]\n    x = imgs.reshape(-1, *imgs.shape[-3:]).float()\n    n, dev = x.shape[0], x.device\n\n    rot = (torch.rand(n, device=dev) - 0.5) * 2 * (AUG_ROT_DEG * np.pi / 180)\n    # Zoom in only. `border` padding repeats the edge row outward, and the edge of this\n    # crop is where the popliteal fossa sits; zooming out would fabricate tissue exactly\n    # where a Baker cyst is looked for.\n    sc = 1.0 + torch.rand(n, device=dev) * AUG_SCALE\n    tx = (torch.rand(n, device=dev) - 0.5) * 2 * AUG_SHIFT\n    ty = (torch.rand(n, device=dev) - 0.5) * 2 * AUG_SHIFT\n    cos, sin = torch.cos(rot) / sc, torch.sin(rot) / sc\n    theta = torch.zeros(n, 2, 3, device=dev, dtype=torch.float32)\n    theta[:, 0, 0], theta[:, 0, 1], theta[:, 0, 2] = cos, -sin, tx\n    theta[:, 1, 0], theta[:, 1, 1], theta[:, 1, 2] = sin, cos, ty\n    grid = F.affine_grid(theta, x.shape, align_corners=False)\n    x = F.grid_sample(x, grid, mode=\"bilinear\", padding_mode=\"border\", align_corners=False)\n\n    scale = 1.0 + (torch.rand(n, 1, 1, 1, device=dev) - 0.5) * 2 * AUG_INTENSITY\n    x = (x * scale).clamp(0, 255)\n    return x.reshape(*lead, *x.shape[-3:]).to(imgs.dtype)\n\n\n@torch.no_grad()\ndef predict(model, cache, mask, idx, dev, img_size=None):\n    \"\"\"Average the logits over the groups of each slot.\n\n    Training sees one group at a time, which acts as augmentation along the stack;\n    inference averages over all of them, so the prediction does not depend on which\n    group a single draw happened to pick. Where the cache holds one group per slot the\n    two coincide.\n    \"\"\"\n    model.eval()\n    out = []\n    for b in range(0, len(idx), EVAL_BATCH):\n        sel = idx[b:b + EVAL_BATCH]\n        m = torch.from_numpy(mask[sel]).to(dev)\n        acc = None\n        for g in range(N_GROUP):\n            # Gathered a group at a time rather than whole and then sliced. The two are\n            # the same pixels, but taking the whole of a study out of the cache allocates\n            # every slice it holds - most of which this pass will not look at until a\n            # later iteration, by which time they have been fetched again. Measured over\n            # a cache of twelve slices, the difference between the two is the difference\n            # between the step being bound by memory and being bound by the encoder.\n            rows = torch.from_numpy(np.ascontiguousarray(\n                cache[sel, :, g * GROUP:(g + 1) * GROUP])).to(dev)\n            with torch.autocast(\"cuda\", enabled=dev.type == \"cuda\"):\n                z = model(rows, m, img_size).float()\n            acc = z if acc is None else acc + z\n        out.append(torch.sigmoid(acc / N_GROUP).cpu().numpy())\n    return np.concatenate(out) if out else np.zeros((0, len(TARGETS)), np.float32)\n\n\ndef macro_auc(y, p):\n    from sklearn.metrics import roc_auc_score\n    return float(np.nanmean([roc_auc_score(y[:, j], p[:, j])\n                             if len(set(y[:, j])) > 1 else np.nan\n                             for j in range(y.shape[1])]))\n"},{"cell_type":"markdown","id":"md-32","metadata":{},"source":"## 7. Validating without fooling yourself\n\nTwo leaks are specific to this setup, and both inflate a validation number without\nimproving anything.\n\n**Shared reports.** Some reports are byte-identical across studies — a template read for an\nunremarkable knee — so every study in such a group receives the same derived target vector,\nand splitting the group across the divide scores the model on a target whose source it was\ntrained on. Studies are therefore assigned by a hash of the report text, which keeps every\nduplicate group whole. One fifth is held out, and the split is fixed rather than rotated.\n\n**Two references, two meanings.** The **holdout** covers a fifth of the corpus, measures\nagreement with the derived targets, and has enough studies per label to separate a real\ndifference from noise — it selects both the epoch within a run and the recipe between runs.\nThe **annotation check** measures agreement with a radiologist's reading of the images,\nwhich is what the competition scores, but only the annotated studies falling in the holdout\ncan be used and there are very few; by the standard-error argument of §2 it is reported and\nnever allowed to arbitrate.\n\nThe annotated studies stay in training, at elevated weight, because they are the only labels\nread from the images rather than from text. That is exactly why the annotation check must be\nrestricted to the holdout: scoring a model on training examples whose answers it saw,\nweighted more heavily than anything else, measures memorisation and reports it as skill.\n"},{"cell_type":"code","id":"code-33","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"def write_submission(pred, studies, test_df, path):\n    \"\"\"Write one submission file from a prediction matrix.\n\n    Predictions are converted to per-column ranks first: the metric reads only order, so\n    ranks discard nothing, and they make files from different configurations directly\n    comparable and safe to average.\n    \"\"\"\n    sub = pd.DataFrame(pd.DataFrame(pred).rank(pct=True).values, columns=TARGETS)\n    sub.insert(0, \"StudyInstanceUID\", studies)\n    sub = test_df[[\"StudyInstanceUID\"]].merge(sub, on=\"StudyInstanceUID\", how=\"left\")\n    sub[TARGETS] = sub[TARGETS].fillna(0.5)\n    sub.to_csv(path, index=False)\n    return sub\n\n\ndef write_benchmark_submission():\n    \"\"\"Write the 0.5 benchmark file immediately.\n\n    A submission that never writes scores nothing at all, which is strictly worse than\n    scoring badly. The try/except around main() covers exceptions, but a kill for memory\n    is a SIGKILL and never reaches it. So a valid file exists from the first second and\n    is overwritten only once real predictions are ready.\n    \"\"\"\n    t = pd.read_csv(ROOT / \"test.csv\")\n    for c in TARGETS:\n        t[c] = 0.5\n    t.to_csv(\"submission.csv\", index=False)\n\n\ndef main():\n    write_benchmark_submission()\n\n    # Weights, if any were attached; otherwise the run learns its own below. Both paths\n    # are kept because the second is what makes this notebook readable on its own - a\n    # fork with nothing attached still trains and still scores - and because the first\n    # cannot be checked by anyone who does not have the package.\n    pkg = find_weights()\n    if pkg is not None:\n        torch.cuda.is_available = lambda: False  # force CPU\n        dev = torch.device(\"cpu\")\n        infer_from_package(pkg, dev)\n        log(\"done\")\n        return\n\n    # Settle where the labels come from before anything expensive runs. The check costs\n    # one CSV header read; discovering the same problem after the cache is built would\n    # cost the whole decode pass, and discovering it never would cost the run.\n    read_labels(pd.read_csv(ROOT / \"train.csv\", usecols=[\"StudyInstanceUID\", \"Report\"]))\n\n    test_df = pd.read_csv(ROOT / \"test.csv\")\n    test_series = pd.read_csv(ROOT / \"test_series.csv\")\n    train_df = pd.read_csv(ROOT / \"train.csv\")\n    train_series = pd.read_csv(ROOT / \"train_series.csv\")\n    log(f\"train {train_df.shape} test {test_df.shape}\")\n\n    both = pd.concat([train_series, test_series])\n    plane_map = dict(zip(both[\"SeriesInstanceUID\"], both[\"Anatomical_Plane\"]))\n\n    log(\"header pass: test\")\n    hte = annotate(walk(\"test_series\"))\n    log(f\"  {len(hte)} test series\")\n    log(\"header pass: train\")\n    htr = annotate(walk(\"train_series\"))\n    log(f\"  {len(htr)} train series\")\n\n    slots_te, slots_tr = pick_slots(hte, plane_map), pick_slots(htr, plane_map)\n    cov = pd.Series([len(v) for v in slots_tr.values()]).describe()\n    log(f\"train slots per study: mean {cov['mean']:.2f} min {cov['min']:.0f} \"\n        f\"max {cov['max']:.0f}\")\n\n    st_tr, Ctr, Mtr = build_cache(slots_tr, plane_map, lat_of(htr, \"train \"), \"train\")\n    st_te, Cte, Mte = build_cache(slots_te, plane_map, lat_of(hte, \"test \"), \"test\")\n\n    # ---- targets ---------------------------------------------------------- #\n    t_lab = time.time()\n    lab = read_labels(train_df)\n    log(f\"derived labels for {len(lab)} studies in {time.time() - t_lab:.1f}s\")\n\n    gold = train_df.set_index(\"StudyInstanceUID\")[TARGETS]\n    gold = gold[gold.notna().all(axis=1)]\n\n    Y = np.zeros((len(st_tr), len(TARGETS)), np.float32)\n    W = np.zeros_like(Y)\n    for i, st in enumerate(st_tr):\n        if st in gold.index:\n            Y[i], W[i] = gold.loc[st].values, 3.0\n        elif st in lab.index:\n            r = lab.loc[st]\n            Y[i] = r[TARGETS].values\n            W[i] = 0.25 + 0.75 * r[[t + \"__conf\" for t in TARGETS]].values\n    keep = np.where(W.sum(1) > 0)[0]\n    log(f\"supervised {len(keep)} of {len(st_tr)} studies (annotated {len(gold)})\")\n\n    # Grouped on report text: some reports are byte-identical across studies and yield\n    # one target vector for all of them, so splitting such a group scores the model on a\n    # target whose source it has already trained on.\n    import hashlib\n    rep = train_df.set_index(\"StudyInstanceUID\")[\"Report\"].fillna(\"\")\n    grp = np.array([int(hashlib.md5(rep.get(s, s).encode()).hexdigest()[:8], 16) % 5\n                    for s in st_tr])\n    va = np.array([i for i in keep if grp[i] == 0])\n    tr = np.array([i for i in keep if grp[i] != 0])\n    if len(va) == 0 or len(tr) < BATCH_STUDIES:\n        cut = max(1, len(keep) // 5)\n        va, tr = keep[:cut], keep[cut:]\n    log(f\"train {len(tr)} / holdout {len(va)} studies\")\n\n    # The annotated studies stay in training - they are the highest-quality labels in\n    # the corpus and there are too few to discard - so the honest annotation check uses\n    # only the ones that fell in the holdout. Evaluating on the rest would be scoring the\n    # model against examples it was trained on, at triple weight, with the true answer.\n    gpos = {s: i for i, s in enumerate(st_tr)}\n    va_set = set(va.tolist())\n    gi = np.array([gpos[s] for s in gold.index if s in gpos and gpos[s] in va_set])\n    gold_y = gold.loc[[st_tr[i] for i in gi]].values.astype(int) if len(gi) else None\n    yv = (Y[va] > 0.5).astype(int)\n    log(f\"annotation check: {len(gi)} of {len(gold)} annotated studies are in the holdout\")\n\n    # ---- fine-tune -------------------------------------------------------- #\n    dev = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n    results, test_preds = {}, {}\n\n    for cfg in RUNS:\n        pitch = CROP_MM / cfg[\"img\"]\n        log(f\"=== {cfg['name']}: {cfg['img']} px, {pitch:.3f} mm/pixel, \"\n            f\"{pitch * 14:.2f} mm per patch token ===\")\n        torch.manual_seed(SEED)\n        model = build_model(UNFREEZE_LAST).to(dev)\n        opt = torch.optim.AdamW([\n            {\"params\": [p for p in model.backbone.parameters() if p.requires_grad],\n             \"lr\": LR_BACKBONE},\n            {\"params\": model.head.parameters(), \"lr\": LR_HEAD},\n        ], weight_decay=WEIGHT_DECAY)\n        steps = max(EPOCHS * (len(tr) // BATCH_STUDIES), 1)\n        sched = torch.optim.lr_scheduler.OneCycleLR(\n            opt, max_lr=[LR_BACKBONE, LR_HEAD], total_steps=steps, pct_start=0.15)\n        scaler = torch.amp.GradScaler(\"cuda\", enabled=dev.type == \"cuda\")\n\n        best, best_state, best_annot = -1.0, None, float(\"nan\")\n        for ep in range(EPOCHS):\n            model.train()\n            perm = np.random.permutation(tr)\n            tot, nstep = 0.0, 0\n            for b in range(0, len(perm) - BATCH_STUDIES + 1, BATCH_STUDIES):\n                sel = perm[b:b + BATCH_STUDIES]\n                rows = torch.from_numpy(Ctr[sel]).to(dev)\n                g = int(torch.randint(N_GROUP, (1,)).item())\n                imgs = augment(take_group(rows, g))\n                m = torch.from_numpy(Mtr[sel]).to(dev)\n                y = torch.from_numpy(Y[sel]).to(dev)\n                w = torch.from_numpy(W[sel]).to(dev)\n                with torch.autocast(\"cuda\", enabled=dev.type == \"cuda\"):\n                    loss = (F.binary_cross_entropy_with_logits(\n                        model(imgs, m, cfg[\"img\"]), y, reduction=\"none\") * w).mean()\n                opt.zero_grad(set_to_none=True)\n                scaler.scale(loss).backward()\n                scaler.step(opt)\n                scaler.update()\n                sched.step()\n                tot += loss.item()\n                nstep += 1\n\n            pv = predict(model, Ctr, Mtr, va, dev, cfg[\"img\"])\n            d = macro_auc(yv, pv)\n            g_auc = float(\"nan\")\n            if gold_y is not None and len(gi):\n                g_auc = macro_auc(gold_y, predict(model, Ctr, Mtr, gi, dev, cfg[\"img\"]))\n            log(f\"  epoch {ep + 1}/{EPOCHS}  loss {tot / max(nstep, 1):.4f}\"\n                f\"  holdout {d:.4f}  annot(n={len(gi)}) {g_auc:.4f}\")\n\n            # Selection reads the holdout alone. The annotation check is reported because\n            # it measures something different - agreement with a reading of the images\n            # rather than of the reports - but only a handful of annotated studies land\n            # in any one holdout, so its sampling error dwarfs the differences between\n            # epochs and it cannot arbitrate between them.\n            if d > best:\n                best, best_annot = d, g_auc\n                best_state = {k: v.detach().cpu().clone()\n                              for k, v in model.state_dict().items()}\n            if time.time() - T0 > TIME_BUDGET:\n                log(\"  time budget reached\")\n                break\n\n        if best_state is not None:\n            model.load_state_dict(best_state)\n        results[cfg[\"name\"]] = (best, best_annot)\n        test_preds[cfg[\"name\"]] = predict(model, Cte, Mte, np.arange(len(st_te)), dev,\n                                          cfg[\"img\"])\n        log(f\"  {cfg['name']}: best holdout {best:.4f} (annot {best_annot:.4f})\")\n        del model, opt, sched, scaler, best_state\n        gc.collect()\n        if dev.type == \"cuda\":\n            torch.cuda.empty_cache()\n\n    log(\"---- summary ----\")\n    for n, (d, g_auc) in results.items():\n        log(f\"  {n:12s} holdout {d:.4f}   annot {g_auc:.4f}\")\n    pick = max(results, key=lambda k: results[k][0])\n    log(f\"best on the holdout: {pick} ({results[pick][0]:.4f})\")\n\n\n    # ---- write every candidate -------------------------------------------- #\n    # One file per configuration, plus the holdout's choice as `submission.csv`. A run\n    # costs a full decode of the corpus whichever configuration wins, so keeping every\n    # arm makes a later change of configuration free rather than another full run.\n    for name, pred in test_preds.items():\n        sub = write_submission(pred, st_te, test_df, f\"submission_{name}.csv\")\n        log(f\"  submission_{name}.csv {sub.shape}; \"\n            f\"nulls {int(sub[TARGETS].isna().sum().sum())}\")\n\n    ens = np.mean([pd.DataFrame(p).rank(pct=True).values for p in test_preds.values()],\n                  axis=0)\n    write_submission(ens, st_te, test_df, \"submission_rankmean.csv\")\n    log(f\"  submission_rankmean.csv (rank mean of {len(test_preds)})\")\n\n    sub = write_submission(test_preds[pick], st_te, test_df, \"submission.csv\")\n    log(f\"submission.csv = {pick}; {sub.shape}; \"\n        f\"nulls {int(sub[TARGETS].isna().sum().sum())}\")\n    print(sub.head().to_string())\n"},{"cell_type":"code","id":"code-34","metadata":{"_kg_hide-input":true,"_kg_hide-output":true},"execution_count":null,"outputs":[],"source":"try:\n    main()\nexcept LabelSourceError:\n    # Deliberately not absorbed: see LabelSourceError. A run that trained on the\n    # wrong labels would finish and write a submission worth submitting by mistake.\n    traceback.print_exc()\n    raise\nexcept Exception:\n    traceback.print_exc()\n    # A submission that fails to write scores nothing at all, so fall back to the\n    # benchmark file rather than dying.\n    t = pd.read_csv(find_root() / \"test.csv\")\n    for c in TARGETS:\n        t[c] = 0.5\n    t.to_csv(\"submission.csv\", index=False)\n    print(\"wrote fallback submission.csv\")\nlog(\"done\")\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.12.0"}},"nbformat":4,"nbformat_minor":5}