{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div style=\"background:linear-gradient(90deg,#1c3f6e 0%,#4d6d9a 45%,#b4636b 78%,#d97e5a 100%);\n            padding:22px 26px;border-radius:10px;color:#fff;font-family:sans-serif\">\n<h1 style=\"margin:0 0 6px 0;font-size:29px\">🦵 Twelve findings from one knee MRI — v4</h1>\n<h3 style=\"margin:0;font-weight:400;opacity:.92\">Graded report targets · physical-scale sampling · adapted DINOv2 · cross-validated ensemble</h3>\n</div>\n\nThe pipeline below is the one described in the sections that follow: a multilingual report\nextractor produces graded targets, series are sampled at a constant physical scale, and a\npartially unfrozen DINOv2 encoder is trained with a per-diagnosis attention head over\nsequence slots.\n\nSix things are different in this version. Each is a change to how the model is trained or\nvalidated, not to what it looks at, and each is behind a switch so it can be measured on\nits own.\n\n| # | Change | Why |\n|---|---|---|\n| 1 | **Cross-validated ensemble** instead of one model on 4/5 of the data | A single fold leaves a fifth of the corpus unused and stakes the submission on one training run. Folds are rank-averaged, which is the only combination the metric reads |\n| 2 | **Honest annotated reference** | The previous selection rule scored the annotated studies *including those it had just trained on*. Here the annotated reference for a fold is the annotated studies held out of that fold, pooled across folds into a complete out-of-fold estimate |\n| 3 | **No vertical flip** | A knee has a fixed superior–inferior orientation. Flipping it swaps the femoral condyle with the tibial plateau, which is exactly the axis three of the OA targets are defined on. Replaced with small rotation, scale and shift, which preserve anatomy |\n| 4 | **Weight EMA** | Epoch selection on a noisy reference is a coin flip between neighbouring epochs. An exponential moving average makes the choice matter less and is usually slightly better than any single epoch |\n| 5 | **Pairwise ranking term** | The metric is a ranking metric; BCE optimises a proxy. A small ranking loss on confident pairs inside each batch optimises it directly |\n| 6 | **Laterality fallback, measured before it is trusted** | The `Laterality` tag is present on about half of studies. The sign of the patient x-coordinate can supply the rest, but only if it agrees with the tag where both exist — so the notebook measures that agreement and enables the fallback only if it passes |\n\nTwo smaller repairs: `read_slot` referenced an undefined `N_SLICE` in its default argument,\nand the header pass now also retains `ImagePositionPatient` and `ImageLaterality`.\n\nEverything else — the extractor, the slot definitions, the physical-scale crop, the cache\nlayout, the head, the encoder schedule — is unchanged, so a comparison against the previous\nversion measures these six things and nothing else.","metadata":{}},{"cell_type":"markdown","source":"## 1. What the score rewards\n\nThe score is the unweighted mean of twelve per-label ROC AUCs:\n\n$$\\text{Score} \\;=\\; \\frac{1}{12}\\sum_{i=0}^{11} \\mathrm{AUC}_i .$$\n\nThree consequences follow directly, and each one removes a design choice.\n\n**Only order matters.** $\\mathrm{AUC}_i$ is invariant under any strictly increasing map\nof the scores for label $i$. Calibration is therefore worth nothing, and a fixed\nthreshold is worth nothing. It also fixes how to combine models: averaging raw\nprobabilities lets whichever model happens to be most confident dominate, whereas\naveraging *ranks* combines the only information the metric reads. Every combination\nbelow is a rank mean.\n\n**Every label costs the same.** Write $M$ for the mean AUC a good model could reach.\nA label left at chance contributes $0.5$ instead of roughly $M$, so it forfeits\n\n$$\\frac{M - 0.5}{12}$$\n\nof the final score no matter how well the other eleven do. At $M = 0.85$ that is\n$0.029$ — larger than the gap between neighbouring places in a mature competition.\nRare findings deserve *more* attention than common ones, not less, because a rare\nfinding is where a model most easily ends up at chance.\n\n**Prevalence drifts are survivable, thresholds are not.** AUC is, in expectation,\ninvariant to the positive rate. The competition states that prevalence is not guaranteed\nto match across the training, public and final sets, which would be fatal for any\naccuracy-like metric and for anything tuned to a threshold. It is not fatal here,\nprovided nothing in the pipeline depends on a cutoff. Nothing below does.\n","metadata":{}},{"cell_type":"markdown","source":"## 2. Where the targets come from\n\nOnly a small subset of the training studies carry the twelve per-condition labels. Every\ntraining study carries the original radiology report, and the data description invites\nderiving labels from it.\n\nThe decisive structural fact is in the schemas rather than in the prose: `train.csv` has\na `Report` column and `test.csv` does not. Text is available when fitting and absent when\npredicting. That rules out a fusion model with a text branch — at inference it would have\nnothing to read — and leaves three admissible uses of the reports:\n\n1. turn them into training targets, then fit a pure imaging model;\n2. use them as an auxiliary training signal, distilled into the image encoder and dropped\n   at inference;\n3. use them to weight studies by how confidently their labels could be read.\n\nThis notebook takes the first and the third. A multilingual rule extractor reads each\nreport clause by clause, deciding for each finding whether the clause asserts it, negates\nit, or hedges it, and emits a score together with a confidence. The confidence becomes a\nsample weight, so a study whose report says nothing about synovitis pulls on the\nsynovitis head far less than one that names it.\n\nTwo details matter more than the extractor's internals.\n\n**Reports are graded, annotations are thresholded.** The reporting radiologist and the\nannotator do not share a threshold. A report that says *small joint effusion* may sit\nagainst a negative annotation, because the annotator marked only effusions they judged\nsignificant. A rule of the form *term present $\\Rightarrow$ positive* is therefore wrong\nby construction. Grading the mention — trace, unqualified, marked — is right, and costs\nnothing, because §1 established that only the order of the scores is read.\n\n**Derived labels are not independent across studies.** A report shared verbatim by\nseveral studies yields one target vector for all of them. That has to be respected when\nsplitting; §7 does.\n","metadata":{}},{"cell_type":"markdown","source":"### Reading a report in nine languages\n\nThe extractor is the first model in this pipeline, so it is built here rather than\nattached as a file. It runs over a few megabytes of text in seconds, and keeping it in\nline means the targets can never be a stale copy of what the current rules would produce.\n\n**Script settles itself; language does not.** Greek and Cyrillic are decided by Unicode\nalone. The Latin-script remainder is scored by stopword counts across the candidate\nlanguages and the winner must beat the runner-up by a margin; what fails that test stays\n`unknown` rather than being guessed. The tempting shortcut — a cascade of substring tests,\n`'the '` for English, `'la '` for French — fails badly here, because `la` is as common in\nSpanish as in French and whichever test runs first swallows both.\n\n**Normalise, then segment, then scope.** Case, diacritics and separators are folded first,\nwhich also repairs a codepoint problem: many Greek reports spell mu with the MICRO SIGN\nU+00B5 rather than U+03BC, and NFKD maps one onto the other. Text is then split into\nclauses, with a heading line attached to the value beneath it, because a report that reads\n`Fractures :` and then `Aucune.` states one thing across two lines and any method that\nsplits them reads a negation as a positive.\n\n**Assertion, negation, hedge.** Within a clause the extractor asks which of three things\nthe sentence is doing. Negation is not an edge case: for several findings most mentions\nare negative, since a report lists what was checked and found intact. Explicit normality\ncounts as negation — *ligamentos cruzados y colaterales dentro de límites normales* is\nevidence of absence, not absence of evidence — except where a tear or a high grade is\nnamed in the same breath.\n\n**Stems, not phrases.** Four targets need an anatomy word and a pathology word together,\nand a lexicon of complete phrases cannot survive morphology: Turkish suffixes possessives\nonto the noun, Croatian and Greek decline it. Matching a stem and then requiring a side\nqualifier within a character window handles inflection without enumerating it, and a\ncharacter window rather than a token window handles word order, which puts the side\nadjective before the noun in English and after it in Greek.\n\n**Grade, do not threshold.** §1 established that only the order of the scores is read, and\n§2 that the annotator's threshold is stricter than the reporting radiologist's. Together\nthose say a mention should be scored by its emphasis — trace, unqualified, marked — and\nnever binarised. Each target also carries a confidence, which becomes the sample weight:\nsilence on a finding is weak evidence, and it should pull on the model weakly.\n\n### How an extractor like this is validated\n\nThis is the part that decides whether any of the above is worth trusting, and it is\nharder than writing the rules, because the obvious measurement is the one that cannot\ncarry the weight.\n\n**The dangerous failure is silent.** A rule that never fires does not raise an error; it\nemits a negative. A lexicon that is complete in English and thin in Greek therefore does\nnot look broken — it looks like a corpus where Greek patients have fewer findings. Worse,\nthe error is not random: language tracks the reporting institution, which tracks the\nscanner and the population, so a gap in one language is a systematic bias aligned with a\nsite rather than noise that averages out.\n\n**Gauge one: agreement, on the annotated subset.** For each target, compare the extracted\nscore against the per-condition annotation and read the AUC. This measures the right\nthing, and it is nearly useless for tuning, because that subset is small. The\nHanley–McNeil approximation for the standard error of an AUC $A$ with $n_p$ positives and\n$n_n$ negatives is\n\n$$\n\\mathrm{SE}(A)=\\sqrt{\\frac{A(1-A)+(n_p-1)(Q_1-A^{2})+(n_n-1)(Q_2-A^{2})}{n_p\\,n_n}},\n\\qquad\nQ_1=\\frac{A}{2-A},\\quad Q_2=\\frac{2A^{2}}{1+A}.\n$$\n\nPut a plausible $A\\approx0.8$ and a rare finding — a handful of positives among a few\ndozen studies — into that expression and the standard error lands near $0.09$, so the 95%\ninterval spans roughly $\\pm0.17$. Competitions are decided by differences an order of\nmagnitude smaller. Choosing between two lexicons on this number is choosing by coin flip,\nand it will feel like signal every time.\n\n**Gauge two: coverage, on the whole corpus.** For each (study, target) pair, record\nwhether any rule fired at all — assertion, negation or hedge. The *silence rate* is the\nfraction where none did. It needs no labels, so it runs on every report rather than on the\nannotated handful, and broken down by language and target it points straight at the\nmissing vocabulary. A common finding that is silent in one language and not another is a\nlexicon gap. A rare finding that is silent nearly everywhere is simply rare, and silence\nthere is correct.\n\nThe two gauges answer different questions and neither substitutes for the other:\n\n| | measures | sample | can decide |\n|---|---|---|---|\n| agreement | is a fired rule *right* | small | whether a target's labels are usable at all |\n| silence rate | does a rule *fire* | whole corpus | which language and which finding to work on next |\n\n**The loop.** Read actual reports in each language before writing any pattern — the\nvocabulary comes from the corpus, not from a translation of the English list. Write rules,\nthen measure both gauges. Then open the disagreements individually and ask what the\nextractor saw, because the aggregate says a target is weak while a handful of cases says\n*why*: a threshold mismatch, a missing negator, a morphological form the lexicon cannot\nreach. Fix, re-measure, and prefer changes that improve coverage on the large gauge over\nchanges that improve agreement on the small one — the first is signal, the second is\nmostly sampling noise.\n","metadata":{}},{"cell_type":"code","source":"from __future__ import annotations\n\nimport re\nimport unicodedata\n\nTARGETS = [\n    \"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\",\n    \"Medial OA\", \"Lateral OA\", \"PF OA\", \"Effusion\",\n    \"Synovitis\", \"Baker's\", \"Contusion\", \"Fracture\",\n]\n\n\n# Turkish dotted/dotless i must be folded before casefolding, otherwise \"İZLENMEZ\"\n# and \"izlenmez\" diverge. ß and the Croatian/Serbian d-with-stroke likewise.\n_PRE = str.maketrans({\n    \"ı\": \"i\", \"İ\": \"i\", \"I\": \"i\", \"ß\": \"ss\", \"đ\": \"d\", \"Đ\": \"d\",\n    \"ø\": \"o\", \"Ø\": \"o\", \"æ\": \"ae\", \"Æ\": \"ae\",\n})\n\n\ndef normalize(text: str) -> str:\n    \"\"\"Fold case, diacritics and separators; keep Greek and Cyrillic letters.\n\n    NFKD decomposition strips Latin accents and Greek tonos alike (ά -> α), which is what\n    we want: reports are inconsistent about accents. It also maps the MICRO SIGN U+00B5\n    to a real mu, which matters because most Greek reports here use the wrong codepoint.\n    \"\"\"\n    if not isinstance(text, str):\n        return \"\"\n    text = text.translate(_PRE).lower()\n    text = unicodedata.normalize(\"NFKD\", text)\n    text = \"\".join(ch for ch in text if not unicodedata.combining(ch))\n    text = text.replace(\"­\", \"\")                    # soft hyphen\n    text = re.sub(r\"[_\\-/\\\\]+\", \" \", text)\n    text = re.sub(r\"[ \\t]+\", \" \", text)\n    return text\n\n\n_SENT_SPLIT = re.compile(r\"(?<=[.;!?])\\s+|\\n+\")\n\n\ndef clauses(text: str):\n    \"\"\"Split into clauses, then attach `header:` lines to the value that follows.\n\n    A report line reading `Fractures :` followed by `Aucune.` is one statement. Splitting\n    on punctuation alone separates the anatomy from its negation and flips the label.\n    \"\"\"\n    norm = normalize(text)\n    raw = [c.strip() for c in _SENT_SPLIT.split(norm) if c and c.strip()]\n\n    merged = []\n    for i, c in enumerate(raw):\n        # A fragment ending in a colon is a heading for the next fragment. Structured\n        # English reports write long ones - \"lateral compartment (meniscus, collateral\n        # ligament complex, cartilage):\" is eight words - so the cap is generous, and\n        # the heading is kept as its own clause too in case it carries the finding.\n        if c.endswith(\":\") and len(c.split()) <= 14 and i + 1 < len(raw):\n            merged.append(c + \" \" + raw[i + 1])\n        merged.append(c)\n    # Comma-separated enumerations inside a long clause hide separate assertions.\n    out = []\n    for c in merged:\n        out.append(c)\n        if len(c.split()) > 25:\n            out.extend(p.strip() for p in c.split(\",\") if len(p.split()) > 2)\n    return out\n\n\ndef _rx(*alts: str) -> re.Pattern:\n    return re.compile(\"|\".join(alts))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"NEGATION = _rx(\n    # en\n    r\"\\bno\\b\", r\"\\bnot\\b\", r\"\\bwithout\\b\", r\"\\bnegative for\\b\", r\"\\babsence\\b\",\n    r\"\\bno evidence\\b\", r\"\\bunremarkable\\b\", r\"\\bfree of\\b\",\n    # es\n    r\"\\bsin\\b\", r\"\\bno hay\\b\", r\"\\bausencia\\b\", r\"\\bausentes?\\b\",\n    # fr\n    r\"\\bpas de\\b\", r\"\\bsans\\b\", r\"\\baucune?\\b\", r\"\\babsence\\b\",\n    # nl\n    r\"\\bgeen\\b\", r\"\\bzonder\\b\", r\"\\bniet\\b\",\n    # de\n    r\"\\bkeine?\\b\", r\"\\bohne\\b\", r\"\\bnicht\\b\",\n    # tr\n    r\"\\byok\\b\", r\"\\byoktur\\b\", r\"izlenmemekte\", r\"saptanmadi\", r\"\\bdegil\\b\",\n    r\"gozlenmemekte\", r\"mevcut degil\", r\"eslik etmiyor\", r\"\\bizlenmedi\\b\",\n    # hr / sr / bs\n    r\"\\bnema\\b\", r\"\\bbez\\b\", r\"\\bnisu\\b\", r\"\\bnije\\b\",\n    # el (accents already stripped)\n    r\"\\bδεν\\b\", r\"\\bχωρις\\b\", r\"ουδεν\",\n    # bg / ru\n    r\"\\bбез\\b\", r\"\\bне\\b\", r\"липсва\", r\"\\bняма\\b\",\n)\n\nNORMALITY = _rx(\n    r\"\\bnormal\", r\"\\bintact\\b\", r\"\\bpreserved\\b\", r\"\\bwithin normal limits\\b\",\n    r\"limites normales\", r\"\\bconservad\", r\"\\bintegr\", r\"\\bnormales\\b\",\n    r\"\\bdoga(l|ll)\\b\", r\"korunmus\", r\"\\bnormaldir\\b\", r\"olagan\",\n    r\"\\buredn\", r\"\\bocuvan\", r\"\\bodrzan\", r\"\\bintakt\",\n    r\"φυσιολογικ\", r\"ακεραι\",\n    r\"unauffallig\", r\"regelrecht\", r\"\\bintakt\\b\",\n    r\"нормал\", r\"запазен\", r\"съхранен\", r\"\\bбез особености\\b\",\n    r\"\\bgaaf\\b\", r\"\\bnormaal\\b\",\n)\n\nUNCERTAIN = _rx(\n    r\"\\bpossible\\b\", r\"\\bprobable\\b\", r\"\\bsuspicious\\b\", r\"\\bsuspected\\b\",\n    r\"cannot (be )?exclude\", r\"\\bmay\\b\", r\"\\bquestionable\\b\", r\"\\bequivocal\\b\",\n    r\"\\bposible\\b\", r\"sin criterios categoricos\", r\"\\bdudos\",\n    r\"\\bmuhtemel\\b\", r\"\\bolasi\\b\", r\"\\bsupheli\\b\", r\"\\bizlenim\",\n    r\"\\bmoguce\\b\", r\"\\bvjerojatno\\b\", r\"\\bsumnja\\b\",\n    r\"πιθαν\", r\"υποπτ\",\n    r\"\\bmoglich\", r\"\\bverdachtig\", r\"\\bfraglich\", r\"\\bV\\.a\\.\\b\",\n    r\"\\bвъзможно\\b\", r\"\\bвероятно\\b\", r\"суспект\",\n    r\"\\bmogelijk\\b\", r\"\\bverdacht\\b\",\n)\n\n# Pathology vocabulary shared by the paired rules.\nTEAR = _rx(\n    r\"\\btear\", r\"\\btorn\\b\", r\"\\brupture\", r\"\\bdisruption\\b\", r\"discontinuit\",\n    r\"\\bavuls\",\n    r\"\\brotura\\b\", r\"\\broturas\\b\", r\"\\bruptura\", r\"\\bdesgarro\", r\"\\broto\\b\",\n    r\"\\bdechirure\", r\"\\bdechire\",\n    r\"\\bscheur\", r\"\\bruptuur\", r\"gescheurd\",\n    r\"\\briss\\b\", r\"einriss\", r\"\\bruptur\", r\"zerreiss\", r\"\\blasion\",\n    r\"\\byirtik\", r\"\\byirtig\", r\"\\bkopma\\b\", r\"butunluk kaybi\", r\"\\brupturu\\b\",\n    r\"\\bpuknuce\", r\"\\bruptur\", r\"\\bprekid\\b\", r\"\\bpukotin\",\n    r\"ρηξη\", r\"ρηξις\", r\"ρηγμα\",\n    r\"руптура\", r\"разкъсв\", r\"разрив\", r\"скъсв\",\n)\n\nDEGEN = _rx(\n    r\"degenerat\", r\"\\bmucoid\\b\", r\"\\bmyxoid\\b\", r\"\\bfray\", r\"\\bfissur\",\n    r\"dejeneratif\", r\"\\bmukoid\\b\", r\"degenerativn\", r\"εκφυλιστ\", r\"дегенерат\",\n    r\"\\bmuco ?ide\\b\", r\"aufgefasert\",\n)\n\nINJURY = _rx(\n    r\"\\binjur\", r\"\\bsprain\", r\"\\blesion\", r\"\\blasion\", r\"\\bedema\\b\", r\"\\boedema\\b\",\n    r\"\\bodem\\b\", r\"\\bedem\\b\", r\"\\bοιδημα\", r\"\\bодем\", r\"\\bедем\", r\"\\bstrain\\b\",\n    r\"\\bhigh signal\\b\", r\"\\bsignal alteration\\b\", r\"\\bhiperintens\", r\"\\bhyperintens\",\n    r\"\\bthicken\", r\"\\bzadebljanje\\b\", r\"\\bverdikking\\b\", r\"\\bdistenzij\",\n    r\"\\blaksite\\b\", r\"\\blaxity\\b\", r\"\\bpartial\\b\", r\"\\bparcijaln\", r\"\\bparcial\",\n    r\"\\bpartiel\", r\"\\bpartiell\",\n)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"ANAT = {\n    \"ACL\": _rx(\n        r\"anterior cruciate\", r\"\\bacl\\b\",\n        r\"cruzado anterior\", r\"\\blca\\b\",\n        r\"croise anterieur\",\n        r\"voorste kruisband\", r\"\\bvkb\\b\",\n        r\"vorderes kreuzband\", r\"vorderen kreuzband\", r\"vordere kreuzband\",\n        r\"on capraz\", r\"\\bocb\\b\",\n        r\"prednji krizni\", r\"prednjeg krizn\",\n        r\"προσθι[οα][^ ]* χιαστ\", r\"προσθιου χιαστου\", r\"χιαστο[^ ]* συνδεσμ\",\n        r\"предна кръстна\", r\"предната кръстна\",\n        # Plural, unqualified: reports routinely clear both cruciates in one clause\n        # (\"Ligamentos cruzados y colaterales dentro de limites normales\"). Without\n        # this, Spanish ACL was silent on 88% of its reports and Dutch on 70%.\n        r\"cruciate ligaments\", r\"ligamentos cruzados\", r\"ligaments croises\",\n        r\"kruisbanden\", r\"kreuzbander\", r\"capraz baglar\", r\"krizn[a-z]* ligament[a-z]*\",\n        r\"χιαστοι συνδεσμ\", r\"χιαστων συνδεσμ\", r\"кръстните връзки\", r\"кръстни връзки\",\n    ),\n    \"MCL\": _rx(\n        r\"medial collateral\", r\"\\bmcl\\b\", r\"tibial collateral\",\n        r\"colateral medial\", r\"colateral interno\", r\"\\blcm\\b\",\n        r\"collateral medial\", r\"collateral interne\",\n        r\"mediale collaterale\", r\"binnenband\",\n        r\"innenband\", r\"mediales? kollateral\",\n        r\"\\bic yan bag\", r\"medial kollateral\", r\"\\biyb\\b\",\n        r\"medijalni kolateraln\", r\"medijalnog kolateraln\",\n        r\"εσω πλαγι\", r\"εσωτερικο πλαγι\",\n        r\"медиален колатерал\", r\"вътрешна странична\",\n        # Same plural pattern as the cruciates.\n        # \"Ligamentos cruzados y colaterales\" separates the noun from its adjective, so\n        # the adjective has to stand alone as a cue.\n        r\"\\bcolaterales\\b\", r\"\\bcollateraux\\b\", r\"\\bcollateralen\\b\", r\"\\bkolateralni\\b\",\n        r\"collateral ligaments\", r\"ligamentos colaterales\", r\"ligaments collateraux\",\n        r\"collaterale banden\", r\"kollateralbander\", r\"seitenbander\", r\"yan baglar\",\n        r\"kolateraln[a-z]* ligament[a-z]*\", r\"πλαγιοι συνδεσμ\", r\"πλαγιων συνδεσμ\",\n        r\"колатерални връзки\", r\"страничните връзки\",\n    ),\n    \"Medial Meniscus\": _rx(\n        r\"medial meniscus\", r\"\\bmm\\b(?= tear)\", r\"medial menisc\",\n        r\"menisco medial\", r\"menisco interno\",\n        r\"menisque medial\", r\"menisque interne\",\n        r\"mediale meniscus\", r\"binnenmeniscus\",\n        r\"innenmeniskus\", r\"medialen? meniskus\", r\"innenmeniskushinterhorn\",\n        r\"medyal menisk\", r\"\\bic menisk\",\n        r\"medijalni meniskus\", r\"medijalnog meniskusa\", r\"medijalnom meniskusu\",\n        r\"εσω μηνισκ\", r\"μηνισκ[^ ]* του εσω\", r\"εσω διαμερισμα[^.]{0,40}μηνισκ\",\n        r\"медиалния менискус\", r\"медиален менискус\", r\"вътрешния менискус\",\n    ),\n    \"Lateral Meniscus\": _rx(\n        r\"lateral meniscus\", r\"lateral menisc\",\n        r\"menisco lateral\", r\"menisco externo\",\n        r\"menisque lateral\", r\"menisque externe\",\n        r\"laterale meniscus\", r\"buitenmeniscus\",\n        r\"aussenmeniskus\", r\"lateralen? meniskus\",\n        r\"lateral menisk\", r\"\\bdis menisk\",\n        r\"lateralni meniskus\", r\"lateralnog meniskusa\", r\"lateralnom meniskusu\",\n        r\"εξω μηνισκ\", r\"μηνισκ[^ ]* του εξω\", r\"εξω διαμερισμα[^.]{0,40}μηνισκ\",\n        r\"латералния менискус\", r\"латерален менискус\", r\"външния менискус\",\n    ),\n}\n\n# Osteoarthritis is rarely written as \"osteoarthritis\". It is written as cartilage loss,\n# chondropathy grade, joint space narrowing, or osteophytes - scoped to a compartment.\nOA_EVIDENCE = _rx(\n    r\"osteoarthrit\", r\"\\barthros\", r\"\\bgonarthros\", r\"\\bosteoarthros\",\n    r\"chondropath\", r\"chondromalac\", r\"condropat\", r\"condromalac\",\n    r\"cartilage loss\", r\"cartilage thinning\", r\"chondral (loss|defect|ulcer|thinning)\",\n    r\"osteophyt\", r\"osteofit\", r\"osteofyt\", r\"osteofito\", r\"osteophyten\",\n    r\"joint space narrowing\", r\"pinzamiento articular\",\n    r\"kikirdak kayb\", r\"kikirdak incelme\", r\"kondropati\", r\"kondral\",\n    r\"kraakbeen(lijden|verlies)\", r\"gonartrose\", r\"artrose\",\n    r\"knorpel(verlust|schaden|defekt)\", r\"arthrose\", r\"gonarthrose\",\n    r\"hrskavic\", r\"hondromalac\", r\"artroz\", r\"osteoartrit\",\n    r\"χονδρ[^ ]*παθ\", r\"αρθριτ\", r\"αρθρωσ\", r\"οστεοφυτ\",\n    r\"αρθρικου χονδρου\", r\"εξαλειψη του αρθρικου χονδρου\",\n    r\"артроз\", r\"хондропат\", r\"остеофит\", r\"хрущял[^.]{0,30}(изтън|увред|дефект)\",\n    r\"ulcera[s]? condral\", r\"cartilago[^.]{0,25}(perdida|adelgaz)\",\n    r\"icrs grade\", r\"outerbridge\",\n)\n\nCOMPARTMENT = {\n    \"Medial OA\": _rx(\n        r\"medial (femorotibial|tibiofemoral|compartment)\",\n        r\"compartimento femorotibial medial\", r\"femorotibial interno\",\n        r\"mediaal femorotibiaal\", r\"mediale femorotibial\",\n        r\"medial femorotibial\", r\"medialen kompartiment\", r\"innere[sn]? kompartiment\",\n        r\"medyal femorotibial\", r\"ic kompartman\", r\"medyal kompartman\",\n        r\"medijaln[^ ]* (femorotibi|odjelj|kompartm)\",\n        r\"εσω διαμερισμα\", r\"εσω κνημιαι\", r\"εσω μηριαι\",\n        r\"медиалн[^ ]* (компартм|отдел|тибиал|феморотиб)\",\n        r\"medial (femoral|tibial) (condyle|plateau)\", r\"condilo femoral medial\",\n        r\"medialen? (femurkondyl|tibiaplateau)\", r\"mediale femorale condyl\",\n    ),\n    \"Lateral OA\": _rx(\n        r\"lateral (femorotibial|tibiofemoral|compartment)\",\n        r\"compartimento femorotibial lateral\", r\"femorotibial externo\",\n        r\"lateraal femorotibiaal\", r\"laterale femorotibial\",\n        r\"lateral femorotibial\", r\"lateralen kompartiment\", r\"aussere[sn]? kompartiment\",\n        r\"lateral femorotibial\", r\"dis kompartman\", r\"lateral kompartman\",\n        r\"lateraln[^ ]* (femorotibi|odjelj|kompartm)\",\n        r\"εξω διαμερισμα\", r\"εξω κνημιαι\", r\"εξω μηριαι\",\n        r\"латералн[^ ]* (компартм|отдел|тибиал|феморотиб)\",\n        r\"lateral (femoral|tibial) (condyle|plateau)\", r\"condilo femoral lateral\",\n        r\"lateralen? (femurkondyl|tibiaplateau)\", r\"laterale femorale condyl\",\n    ),\n    \"PF OA\": _rx(\n        r\"patellofemoral\", r\"femoropatellar\", r\"femoropatelar\", r\"patelofemoral\",\n        r\"retropatellar\", r\"retrorotulian\", r\"\\btrochlea\", r\"\\btroclea\", r\"\\btroklea\",\n        r\"\\bpatella\\b\", r\"\\bpatellar\\b\", r\"\\brotulian\", r\"\\brotula\\b\", r\"\\bpatele\\b\",\n        r\"\\bpatellae?\\b\", r\"patellofemoraal\", r\"femoropatellair\",\n        r\"επιγονατιδ\", r\"μηροεπιγονατιδ\", r\"τροχιλ\",\n        r\"пател\", r\"феморопател\", r\"тролх\",\n        r\"anterior compartment\", r\"compartimento anterior\", r\"prednj[^ ]* odjeljk\",\n    ),\n}\n\n# Self-declaring findings: the term itself is the finding.\nDIRECT = {\n    \"Effusion\": _rx(\n        r\"\\beffusion\", r\"joint fluid\", r\"intra ?articular fluid\", r\"\\bhydrops\\b\",\n        r\"derrame articular\", r\"\\bderrame\\b\", r\"liquido articular\",\n        r\"epanchement\",\n        r\"gewrichtsvocht\", r\"\\bvocht\\b\", r\"\\bhydrops\\b\", r\"gewrichtseffusie\",\n        r\"gelenkerguss\", r\"\\berguss\\b\", r\"gelenksergu\",\n        # \"diz eklemi ici sivi miktari ... artmis\" and \"eklem icerisinde yaygin sivi\n        # artisi\" both occur; the noun takes a possessive suffix, so `eklem ` alone\n        # misses. Match the stem plus any suffix.\n        r\"eklem\\w* ic\\w* sivi\", r\"efuzyon\", r\"eklem sivisi\",\n        r\"sivi (miktari|artisi|birikimi)\", r\"sivi artis\", r\"\\bsivi\\b[^.]{0,25}artmis\",\n        r\"\\bizljev\", r\"\\bizliv\", r\"zglobn[^ ]* tekucin\", r\"\\bhidrops\\b\",\n        r\"αρθρικ[^ ]* υγρ\", r\"υγρου ενδαρθρικα\", r\"ενδαρθρικ[^ ]* υγρ\", r\"ποσοτητα υγρου\",\n        r\"ενδαρθρικ\", r\"αρθρικη συλλογη\", r\"υγρο στην αρθρωση\", r\"υγρου στην αρθρωση\",\n        r\"ставен излив\", r\"излив\", r\"ставна течност\", r\"синовиална течност\",\n    ),\n    \"Synovitis\": _rx(\n        r\"synovit\", r\"sinovit\", r\"synovial (thickening|proliferation|hypertroph)\",\n        r\"synovitis\", r\"synoviale? (verdikking|proliferatie)\",\n        r\"synovialitis\", r\"synovialis(verdickung|proliferation)\",\n        r\"sinovijalitis\", r\"sinovitis\", r\"zadebljanje sinovij\",\n        r\"υμενιτιδα\", r\"συνοβιτιδα\", r\"υμενικ[^ ]* υπερτροφ\", r\"αρθρικου υμεν\",\n        r\"синовит\", r\"синовиал[^ ]* (задебел|пролифер)\",\n        r\"verdikkingen van (het )?synovium\", r\"pannus\",\n    ),\n    \"Baker's\": _rx(\n        r\"baker\", r\"popliteal cyst\", r\"quiste popliteo\", r\"quistes popliteos\",\n        r\"kyste poplite\", r\"popliteale? cyst\", r\"poplitealzyste\", r\"bakerzyste\",\n        r\"popliteal kist\", r\"\\bbakerova\\b\", r\"poplitealn[^ ]* cist\",\n        r\"κυστη baker\", r\"πολυχωρη συνοβιακη κυστη\", r\"κυστη του baker\",\n        r\"киста на бейкър\", r\"бейкърова киста\", r\"поплитеална киста\",\n        r\"gastrocnemio ?semimembranos\", r\"gastrocnemius semimembranosus burs\",\n    ),\n    \"Contusion\": _rx(\n        r\"\\bcontusion\", r\"bone bruise\", r\"bone marrow (o?edema|contusion)\",\n        r\"\\bkontuz\", r\"medular bone o?edema\", r\"marrow o?edema\",\n        r\"contusion osea\", r\"edema oseo\", r\"edema de medula osea\",\n        r\"oedeme osseux\", r\"contusion osseuse\",\n        r\"botcontusie\", r\"botoedeem\", r\"beenmergoedeem\", r\"botmergoedeem\",\n        r\"knochenmarkodem\", r\"knochenodem\", r\"kontusion\", r\"bone bruise\",\n        r\"kemik kontuzyonu\", r\"kemik iligi odemi\", r\"kemik odemi\",\n        r\"kostani edem\", r\"edem kosti\", r\"kontuzij\",\n        r\"οστεομυελικ[^ ]* οιδημα\", r\"οστικο οιδημα\", r\"μυελικο οιδημα\",\n        r\"костномозъчен едем\", r\"костен едем\", r\"контузионен\",\n    ),\n    \"Fracture\": _rx(\n        r\"\\bfractur\", r\"\\bfract\\b\",\n        r\"\\bfractura\", r\"\\bfracturas\\b\",\n        r\"\\bfractuur\", r\"\\bbreuk\\b\",\n        r\"\\bfraktur\", r\"\\bbruch\\b\",\n        r\"\\bkirik\\b\", r\"\\bkirigi\\b\", r\"\\bkirik\\b\",\n        r\"\\bfraktur\", r\"\\bprijelom\", r\"impresijsk[^ ]* fraktur\",\n        r\"καταγμα\", r\"καταγματ\",\n        r\"фрактур\", r\"счупван\", r\"фисур\",\n        r\"insufficiency fracture\", r\"stress fracture\", r\"avulsion fracture\",\n        r\"subchondral fracture\", r\"subkondral kiri\",\n    ),\n}\n\n# Terms that look like a finding but are not the finding being scored.\nDECOY = {\n    \"Fracture\": _rx(r\"no fracture\", r\"microfractur\", r\"\\bfracture (risk|prophyla)\"),\n    \"Baker's\": _rx(r\"meniscal cyst\", r\"quiste meniscal\", r\"ganglion\"),\n}\n\nPAIRED = {\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\"}\nOA_TARGETS = {\"Medial OA\", \"Lateral OA\", \"PF OA\"}\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"STEM_MENISCUS = _rx(r\"menisc\\w*\", r\"menisk\\w*\", r\"μηνισκ\\w*\", r\"мениск\\w*\")\nSTEM_CRUCIATE = _rx(r\"cruciate\", r\"cruzado\", r\"croise\", r\"kruisband\", r\"kreuzband\",\n                    r\"capraz bag\\w*\", r\"krizn\\w*\", r\"χιαστ\\w*\", r\"кръстн\\w*\",\n                    r\"\\bacl\\b\", r\"\\bpcl\\b\", r\"\\blca\\b\", r\"\\blcp\\b\", r\"\\bvkb\\b\",\n                    r\"\\bhkb\\b\", r\"\\bocb\\b\", r\"\\bacb\\b\")\nSTEM_COLLATERAL = _rx(r\"collateral\\w*\", r\"colateral\\w*\", r\"kollateral\\w*\",\n                      r\"collaterale\\w*\", r\"kolateraln\\w*\", r\"yan bag\\w*\",\n                      r\"πλαγι\\w*\", r\"колатерал\\w*\", r\"странич\\w*\",\n                      r\"innenband\\w*\", r\"aussenband\\w*\", r\"binnenband\\w*\",\n                      r\"\\bmcl\\b\", r\"\\blcl\\b\", r\"\\blcm\\b\", r\"\\biyb\\b\")\n\nSIDE_MEDIAL = _rx(r\"\\bmedial\\w*\", r\"\\bmedyal\\w*\", r\"\\bmedijaln\\w*\", r\"\\bmediaal\\w*\",\n                  r\"\\bmediale\\w*\", r\"\\binterno\\w*\", r\"\\binterne\\w*\", r\"\\binnen\\w*\",\n                  r\"\\bic\\b\", r\"\\bunutarnj\\w*\", r\"\\bεσω\\w*\", r\"\\bεσωτερικ\\w*\",\n                  r\"\\bмедиал\\w*\", r\"\\bвътреш\\w*\", r\"\\btibial collateral\\b\",\n                  r\"\\bbinnen\\w*\", r\"\\bmediaal\\b\")\nSIDE_LATERAL = _rx(r\"\\blateral\\w*\", r\"\\bexterno\\w*\", r\"\\bexterne\\w*\", r\"\\bdis\\b\",\n                   r\"\\blateraln\\w*\", r\"\\baussen\\w*\", r\"\\bbuiten\\w*\", r\"\\bεξω\\w*\",\n                   r\"\\bεξωτερικ\\w*\", r\"\\bлатерал\\w*\", r\"\\bвъншн\\w*\",\n                   r\"\\bfibular collateral\\b\", r\"\\bvanjsk\\w*\")\nSIDE_ANTERIOR = _rx(r\"\\banterior\\w*\", r\"\\bant\\b\", r\"\\bon\\b\", r\"\\bprednj\\w*\",\n                    r\"\\bvorder\\w*\", r\"\\bvoorste\\b\", r\"\\bπροσθι\\w*\", r\"\\bпредн\\w*\",\n                    r\"\\banteriyor\\w*\", r\"\\bavant\\b\", r\"\\bant[eé]rieur\\w*\")\n\n# Fracture is the target whose stem varies most across the corpus.\nSTEM_FRACTURE = _rx(r\"fractur\\w*\", r\"fraktur\\w*\", r\"fractuur\\w*\", r\"\\bfract\\b\",\n                    r\"kiri[kgğ]\\w*\", r\"prijelom\\w*\", r\"lom kosti\", r\"\\bbreuk\\w*\",\n                    r\"\\bbruch\\w*\", r\"καταγμα\\w*\", r\"καταγματ\\w*\", r\"фрактур\\w*\",\n                    # NOT a bare `fissur\\w*`: \"fisuras condrales\" and \"full thickness\n                    # fissures in the articular cartilage\" describe cartilage, not bone.\n                    # The stem has to be anchored to a bone word to mean a fracture.\n                    r\"счупван\\w*\", r\"fisur\\w* (osea|oseas|kost)\", r\"fissur\\w* kost\")\n\nSTEM_OA_COMPARTMENT = _rx(r\"compartment\\w*\", r\"compartimento\\w*\", r\"compartiment\\w*\",\n                          r\"kompartman\\w*\", r\"kompartiment\\w*\", r\"odjelj\\w*\",\n                          r\"διαμερισμα\\w*\", r\"компартм\\w*\", r\"\\bотдел\\w*\",\n                          r\"femorotibial\\w*\", r\"femorotibiaal\\w*\", r\"tibiofemoral\\w*\",\n                          r\"femoro tibial\\w*\", r\"κνημιαι\\w*\", r\"μηριαι\\w*\",\n                          r\"femoral condyl\\w*\", r\"tibial plateau\\w*\",\n                          r\"condilo femoral\", r\"platillo tibial\", r\"tibiaplateau\\w*\",\n                          r\"femurkondyl\\w*\", r\"femoralne? kondil\\w*\",\n                          r\"tibijaln\\w* plato\", r\"femoral kondil\\w*\",\n                          r\"tibia plato\", r\"tibyal plato\")\n\n\ndef _near(clause: str, stem_rx: re.Pattern, qual_rx: re.Pattern, window: int = 55):\n    \"\"\"True if a stem match has a qualifier within `window` characters either side.\n\n    Character windows rather than token windows, because word order differs: English\n    puts the side before the noun, Greek and Bulgarian often after, and Turkish\n    attaches it as a separate preceding adjective.\n    \"\"\"\n    for m in stem_rx.finditer(clause):\n        lo = max(0, m.start() - window)\n        hi = min(len(clause), m.end() + window)\n        if qual_rx.search(clause[lo:hi]):\n            return True\n    return False\n\n\n# concept -> (stem, side) pairs used in addition to the phrase lexicons above\nSTEM_RULES = {\n    \"ACL\": (STEM_CRUCIATE, SIDE_ANTERIOR),\n    \"MCL\": (STEM_COLLATERAL, SIDE_MEDIAL),\n    \"Medial Meniscus\": (STEM_MENISCUS, SIDE_MEDIAL),\n    \"Lateral Meniscus\": (STEM_MENISCUS, SIDE_LATERAL),\n    \"Medial OA\": (STEM_OA_COMPARTMENT, SIDE_MEDIAL),\n    \"Lateral OA\": (STEM_OA_COMPARTMENT, SIDE_LATERAL),\n}\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"SEV_LOW = _rx(\n    r\"\\bsmall\\b\", r\"\\bminimal\\b\", r\"\\btrace\\b\", r\"\\bmild\\b\", r\"\\bslight\\b\",\n    r\"\\btiny\\b\", r\"\\bscant\\b\", r\"\\bmimimal\\b\", r\"\\bdiscrete\\b\", r\"\\bfocal\\b\",\n    r\"\\bleve\\b\", r\"\\bminim\", r\"\\bpeque\", r\"\\bligero\\b\", r\"\\bescaso\\b\", r\"\\bdiscreto\\b\",\n    r\"\\bhafif\\b\", r\"\\bminimal\\b\", r\"\\baz miktarda\\b\", r\"\\bsilik\\b\",\n    r\"\\bmanja\\b\", r\"\\bmanji\\b\", r\"\\bblago\\b\", r\"\\bdiskretn\", r\"\\bmalo\\b\",\n    r\"\\bgering\", r\"\\bdiskret\", r\"\\bkleine?r?\\b\", r\"\\bwenig\\b\", r\"\\bzarte?\\b\",\n    r\"\\bbeperkte?\\b\", r\"\\bgeringe\\b\", r\"\\bweinig\\b\", r\"\\blichte?\\b\",\n    r\"\\bηπι\", r\"\\bμικρ\", r\"\\bελαχιστ\",\n    r\"\\bминимал\", r\"\\bлек\", r\"\\bмалк\", r\"\\bнеголям\",\n)\n\nSEV_HIGH = _rx(\n    r\"\\blarge\\b\", r\"\\bmarked\\b\", r\"\\bmassive\\b\", r\"\\bsevere\\b\", r\"\\bextensive\\b\",\n    r\"\\bmoderate\\b\", r\"\\bgross\\b\", r\"\\bsignificant\\b\", r\"\\babundant\\b\", r\"\\btense\\b\",\n    r\"\\bmoderad\", r\"\\bimportante\\b\", r\"\\bsevera?\\b\", r\"\\bmarcad\", r\"\\bcuantios\",\n    r\"\\bbelirgin\\b\", r\"\\byaygin\\b\", r\"\\bileri\\b\", r\"\\bciddi\\b\", r\"\\bbol\\b\",\n    r\"\\bopsezan\\b\", r\"\\bveliki\\b\", r\"\\bizrazit\", r\"\\bznacajn\", r\"\\bumjeren\",\n    r\"\\bausgepragt\", r\"\\bdeutlich\", r\"\\bmassiv\", r\"\\bmassig\", r\"\\bgross\",\n    r\"\\buitgebreid\", r\"\\bgevorderd\", r\"\\bveel\\b\", r\"\\bmatige?\\b\",\n    r\"\\bμετρι\", r\"\\bμεγαλ\", r\"\\bεκτεταμεν\", r\"\\bευμεγεθ\", r\"\\bσοβαρ\",\n    r\"\\bголям\", r\"\\bизразен\", r\"\\bзначим\", r\"\\bумерен\", r\"\\bобилен\",\n)\n\n# OA is often asserted for the whole joint rather than per compartment\n# (\"tricompartmental osteoarthritis\", \"gonarthrose\", \"incipient OA of all three\n# compartments\"). Those statements are evidence for all three OA targets.\nGLOBAL_OA = _rx(\n    r\"tri ?compartment\", r\"all three compartment\", r\"global(ised)? (oa|osteoarthrit)\",\n    r\"\\bgonarthros\", r\"\\bgonartros\", r\"\\bgonarthrose\", r\"\\bgonartrose\",\n    r\"osteoarthritis of the knee\", r\"artrosis (de |)(la )?rodilla\", r\"knee osteoarthrit\",\n    r\"\\bdiz osteoartrit\", r\"\\bgonartroz\", r\"artroza koljena\",\n    r\"οστεοαρθριτιδα\", r\"αρθριτιδα του γονατος\",\n    r\"артроза на колянната\", r\"гонартроз\",\n    r\"degenerative joint disease\", r\"\\bdjd\\b\",\n)\n\n# A bare \"bone marrow oedema\" is not a contusion when it sits under a cartilage\n# defect: subchondral oedema beneath a worn compartment is reactive degenerative signal,\n# and reading it as a bruise turns every osteoarthritic knee into a trauma case.\nDEGENERATIVE_MARROW = _rx(\n    r\"subchondral\", r\"subcondral\", r\"subkondral\", r\"supkondraln\", r\"subchondraln\",\n    r\"υποχονδρι\", r\"субхондрал\", r\"subchondrale?\",\n    r\"\\bcyst\", r\"\\bquist\", r\"\\bzyste\\b\", r\"\\bcistic\", r\"reactive\", r\"reactivo\",\n)\n\nTRAUMA = _rx(\n    r\"\\bbruise\\b\", r\"\\bcontusion\", r\"\\bkontuz\", r\"\\bcontusion osea\\b\",\n    r\"\\btrauma\", r\"\\bimpaction\\b\", r\"\\bpivot shift\\b\", r\"\\bkissing\\b\",\n    r\"\\bacute\\b\", r\"\\bagudo\\b\", r\"\\bakut\", r\"\\bpivot kaymasi\\b\",\n    r\"\\bcontusion osseuse\\b\", r\"\\bbone bruise\\b\", r\"\\bbotcontusie\\b\",\n    r\"\\bконтузион\", r\"\\bμωλωπ\", r\"\\bkontuzij\",\n)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def _polarity(clause: str, anchor_end: int) -> str:\n    \"\"\"Classify one clause as positive, negative or uncertain for a matched term.\n\n    Scope is the whole clause. Clause segmentation already keeps statements short, and\n    a window in characters mis-scopes badly across languages with different word orders -\n    Turkish puts its negator at the end of the sentence, English at the front.\n    \"\"\"\n    if UNCERTAIN.search(clause):\n        return \"uncertain\"\n    if NEGATION.search(clause):\n        return \"negative\"\n    if NORMALITY.search(clause):\n        # \"meniscus normal\" negates; \"normal ... but tear\" does not.\n        if TEAR.search(clause) or re.search(r\"\\bgrade [34]\\b\", clause):\n            return \"positive\"\n        return \"negative\"\n    return \"positive\"\n\n\nclass _Matcher:\n    \"\"\"Phrase lexicon first, stem+side proximity as the fallback.\n\n    Exposes `.search` so it drops into the same slot as a compiled pattern.\n    \"\"\"\n\n    def __init__(self, phrase_rx, stem=None, side=None, window=55):\n        self.phrase_rx = phrase_rx\n        self.stem = stem\n        self.side = side\n        self.window = window\n\n    def search(self, clause):\n        m = self.phrase_rx.search(clause)\n        if m is not None:\n            return m\n        if self.stem is not None and _near(clause, self.stem, self.side, self.window):\n            return self.stem.search(clause)\n        return None\n\n\nANAT_MATCH = {\n    tgt: _Matcher(ANAT[tgt], *STEM_RULES[tgt]) for tgt in PAIRED\n}\nCOMPARTMENT_MATCH = {\n    \"Medial OA\": _Matcher(COMPARTMENT[\"Medial OA\"], *STEM_RULES[\"Medial OA\"]),\n    \"Lateral OA\": _Matcher(COMPARTMENT[\"Lateral OA\"], *STEM_RULES[\"Lateral OA\"]),\n    \"PF OA\": _Matcher(COMPARTMENT[\"PF OA\"]),\n}\nDIRECT_MATCH = {\n    tgt: _Matcher(_rx(rx.pattern, STEM_FRACTURE.pattern) if tgt == \"Fracture\" else rx)\n    for tgt, rx in DIRECT.items()\n}\n\n\ndef _severity(clause: str) -> float:\n    \"\"\"Weight one positive mention by how emphatic the sentence is.\n\n    Ordered, not calibrated. A \"moderate effusion\" must outrank a \"trace effusion\" and\n    both must outrank silence; the absolute numbers do not matter to AUC.\n    \"\"\"\n    high = SEV_HIGH.search(clause) is not None\n    low = SEV_LOW.search(clause) is not None\n    if high and not low:\n        return 1.0\n    if low and not high:\n        return 0.45\n    return 0.75                       # unqualified mention\n\n\ndef _is_structure_heading(clause: str) -> bool:\n    \"\"\"A short line ending in a colon names a section; it asserts nothing by itself.\n\n    clauses() emits both the heading joined to the value beneath it and the heading on its\n    own. For the five targets whose anatomy word *is* the finding - effusion, synovitis,\n    Baker's, contusion, fracture - the bare heading matches the term with no negation in\n    scope, so a structured report reading `Fractures :` / `Aucune.` scored a confident\n    positive off the heading while the joined clause correctly read the negation. Short\n    headings are therefore skipped as assertions; the joined clause carries the meaning.\n    \"\"\"\n    return clause.endswith(\":\") and len(clause.split()) <= 5\n\n\ndef _score_clauses(cls, anat_rx, path_rx=None, decoy_rx=None, context_penalty=None,\n                   context_bonus=None):\n    \"\"\"Accumulate graded evidence over clauses for one target.\n\n    Returns (score, confidence, n_pos, n_neg). Positives are graded by severity and by\n    optional context regexes; negatives only matter when nothing positive was found,\n    because reports assert normality for every structure they check.\n    \"\"\"\n    n_pos = n_neg = n_unc = 0\n    best = 0.0\n    for c in cls:\n        m = anat_rx.search(c)\n        if not m:\n            continue\n        if decoy_rx is not None and decoy_rx.search(c):\n            continue\n        if path_rx is not None and not path_rx.search(c):\n            if NORMALITY.search(c) and not NEGATION.search(c):\n                n_neg += 1\n            continue\n        pol = _polarity(c, m.end())\n        if pol == \"positive\" and _is_structure_heading(c):\n            continue\n        if pol == \"positive\":\n            n_pos += 1\n            w = _severity(c)\n            if context_penalty is not None and context_penalty.search(c):\n                w *= 0.45\n            if context_bonus is not None and context_bonus.search(c):\n                w = min(1.0, w * 1.35)\n            best = max(best, w)\n        elif pol == \"negative\":\n            n_neg += 1\n        else:\n            n_unc += 1\n            best = max(best, 0.30)\n\n    if n_pos or n_unc:\n        # 0.52 .. 0.95, ordered by the strongest single mention, nudged by repetition.\n        score = min(0.95, 0.50 + 0.42 * best + 0.03 * min(n_pos, 3))\n        conf = min(1.0, 0.55 + 0.15 * n_pos)\n    elif n_neg:\n        score = max(0.04, 0.20 - 0.04 * n_neg)\n        conf = min(0.9, 0.45 + 0.12 * n_neg)\n    else:\n        score, conf = 0.28, 0.05          # silence sits above asserted-negative\n    return score, conf, n_pos, n_neg\n\n\ndef extract(report: str) -> dict:\n    \"\"\"Extract twelve (score, confidence) pairs from one report.\"\"\"\n    cls = clauses(report)\n    out = {}\n    path_paired = _rx(TEAR.pattern, DEGEN.pattern, INJURY.pattern)\n\n    for tgt in TARGETS:\n        if tgt in PAIRED:\n            s, c, npos, nneg = _score_clauses(cls, ANAT_MATCH[tgt], path_paired)\n        elif tgt in OA_TARGETS:\n            s, c, npos, nneg = _score_clauses(cls, COMPARTMENT_MATCH[tgt], OA_EVIDENCE)\n        elif tgt == \"Contusion\":\n            # Reactive subchondral oedema under a cartilage defect is osteoarthritis,\n            # not a bruise. Explicit trauma wording pushes the other way.\n            s, c, npos, nneg = _score_clauses(cls, DIRECT_MATCH[tgt], None, DECOY.get(tgt),\n                                              context_penalty=DEGENERATIVE_MARROW,\n                                              context_bonus=TRAUMA)\n        else:\n            s, c, npos, nneg = _score_clauses(cls, DIRECT_MATCH[tgt], None, DECOY.get(tgt))\n        out[tgt] = s\n        out[tgt + \"__conf\"] = c\n        out[tgt + \"__npos\"] = npos\n        out[tgt + \"__nneg\"] = nneg\n\n    # --- cross-target corrections ------------------------------------------ #\n    # A whole-joint osteoarthritis statement is evidence for every compartment that was\n    # not separately assessed. Without this, \"incipient OA of all three compartments\"\n    # scores zero on all three OA targets.\n    g_hits = [c for c in cls if GLOBAL_OA.search(c) and _polarity(c, 0) == \"positive\"]\n    if g_hits:\n        gscore = 0.50 + 0.42 * max(_severity(c) for c in g_hits)\n        for tgt in OA_TARGETS:\n            if out[tgt + \"__npos\"] == 0 and out[tgt + \"__nneg\"] == 0:\n                out[tgt] = max(out[tgt], gscore * 0.92)\n                out[tgt + \"__conf\"] = max(out[tgt + \"__conf\"], 0.4)\n\n    # Synovitis is frequently visible on the images and absent from the text, so silence\n    # is weak evidence of absence here in a way it is not for other findings. Effusion is\n    # its most reliable textual proxy - the two share a mechanism - so a silent synovitis\n    # inherits a fraction of the effusion evidence instead of falling to the floor.\n    if out[\"Synovitis__npos\"] == 0 and out[\"Synovitis__nneg\"] == 0:\n        out[\"Synovitis\"] = max(out[\"Synovitis\"], 0.28 + 0.45 * (out[\"Effusion\"] - 0.28))\n\n    return out\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Three repairs to the extractor\n\nThese were found by running the section above against synthetic reports in the languages it\nclaims to cover. All three are silent failures of exactly the kind §2 warns about: a rule\nthat stops firing emits a negative and raises nothing.\n\n**A bare heading asserted a finding.** `clauses()` emits a heading joined to the value\nbeneath it *and* the heading on its own. For the five targets whose anatomy word is itself\nthe finding, the lone heading matched the term with no negation in scope, so a structured\nreport reading `Fractures :` / `Aucune.` scored 0.84 for fracture off the heading while the\njoined clause correctly read the negation. Short headings no longer assert; the joined\nclause still carries the meaning, so `Fractures :` / `Fracture non deplacee du plateau\ntibial.` still scores 0.88.\n\n**`no fracture` sat in the decoy list.** A decoy skips the clause entirely, so the single\nmost common English negative statement about fractures produced silence rather than an\nasserted negative — the study then pulled on the fracture head with the weight of a report\nthat never mentioned fractures at all. `microfractur` and `fracture risk` stay: those are a\nsurgical procedure and a prediction, neither of which is a fracture.\n\n**`none` and `nil` were not negators.** `Effusion:` / `None.` read as a positive effusion.\nThe additions below are deliberately conservative — each was checked against a real finding\nthat must survive it. One candidate was rejected on that test: French `\\bnon\\b` would have\nnegated `fracture non deplacee`, which is a displaced-fracture qualifier, not a negation.","metadata":{}},{"cell_type":"code","source":"# Additions are unioned onto the compiled patterns rather than edited into them, so the\n# diff against the original lexicon stays readable and every entry can be reverted alone.\n\nNEGATION = _rx(\n    NEGATION.pattern,\n    r\"\\bnone\\b\", r\"\\bnil\\b\",                       # \"Effusion: none.\"\n    r\"not (identified|seen|visuali[sz]ed|demonstrated|detected|present|appreciated)\",\n    r\"\\bnegative\\b\",\n    r\"\\bningun[ao]?\\b\", r\"no se (identifica|aprecia|visualiza|reconoce)\",\n    r\"\\bnenhum\\w*\",\n    r\"niet zichtbaar\", r\"\\bafwezigheid\\b\",\n    r\"izlenmemistir\", r\"gorulmemistir\",\n    # Rejected: r\"\\bnon\\b\". It negates \"fracture non deplacee\", which is a\n    # non-displaced fracture - a positive finding with a qualifier, not an absence.\n)\n\nNORMALITY = _rx(\n    NORMALITY.pattern,\n    r\"\\bwnl\\b\", r\"sans particularite\", r\"sin particularidades\",\n)\n\n# A decoy skips the clause; a negation scores it. \"no fracture\" belongs to the second.\nDECOY[\"Fracture\"] = _rx(r\"microfractur\", r\"\\bfracture (risk|prophyla)\")\n\n_checks = {\n    (\"Fracture\", \"Fractures :\\nAucune.\"): \"lt\",\n    (\"Fracture\", \"No fracture is identified.\"): \"lt\",\n    (\"Fracture\", \"Fractures :\\nFracture non deplacee du plateau tibial.\"): \"gt\",\n    (\"Fracture\", \"Microfracture of the trochlea was performed.\"): \"silent\",\n    (\"Effusion\", \"Effusion:\\nNone.\"): \"lt\",\n    (\"Effusion\", \"Joint effusion: nil.\"): \"lt\",\n    (\"Effusion\", \"Moderate joint effusion.\"): \"gt\",\n    (\"Fracture\", \"Fracturas :\\nNinguna.\"): \"lt\",\n    (\"Baker's\", \"Quiste de Baker :\\nPresente.\"): \"gt\",\n}\nfor (tgt, text), want in _checks.items():\n    v = extract(text)[tgt]\n    ok = (v < 0.3) if want == \"lt\" else (v > 0.5) if want == \"gt\" else (0.25 < v < 0.32)\n    assert ok, f\"{want}: {tgt} scored {v:.2f} on {text!r}\"\nprint(\"lexicon repairs verified on 9 regression cases\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- extractor self-test -------------------------------------------------------------\n# Every claim the section above makes about the extractor is checked here on a synthetic\n# report in the relevant language. Rerun after any lexicon edit: a rule that stops firing\n# does not raise, it silently emits a negative.\n\ndef _s(text):\n    return extract(text)\n\n_cases = {\n    \"en_tear\":     \"Complete tear of the anterior cruciate ligament.\",\n    \"en_intact\":   \"The anterior cruciate ligament is intact.\",\n    \"es_group\":    \"Ligamentos cruzados y colaterales dentro de limites normales.\",\n    \"fr_heading\":  \"Fractures :\\nAucune.\",\n    \"tr_effusion\": \"Eklem icerisinde yaygin sivi artisi mevcuttur.\",\n    \"el_meniscus\": \"Ρηξη του εσω μηνισκου.\",\n    \"eff_large\":   \"Large joint effusion.\",\n    \"eff_trace\":   \"Trace joint effusion.\",\n    \"eff_silent\":  \"The menisci are intact.\",\n    \"global_oa\":   \"Tricompartmental osteoarthritis.\",\n    \"degen_marrow\": \"Subchondral bone marrow oedema beneath a full thickness cartilage defect.\",\n    \"true_bruise\": \"Bone bruise of the lateral femoral condyle after acute trauma.\",\n}\nR = {k: _s(v) for k, v in _cases.items()}\n\nassert R[\"en_tear\"][\"ACL\"] > 0.5, R[\"en_tear\"][\"ACL\"]\nassert R[\"en_intact\"][\"ACL\"] < 0.3, R[\"en_intact\"][\"ACL\"]\n# plural group statement clears both cruciates and both collaterals at once\nassert R[\"es_group\"][\"ACL\"] < 0.3 and R[\"es_group\"][\"MCL\"] < 0.3, (\n    R[\"es_group\"][\"ACL\"], R[\"es_group\"][\"MCL\"])\n# heading on one line, value on the next: the merge in clauses() is what keeps this negative\nassert R[\"fr_heading\"][\"Fracture\"] < 0.3, R[\"fr_heading\"][\"Fracture\"]\nassert R[\"tr_effusion\"][\"Effusion\"] > 0.5, R[\"tr_effusion\"][\"Effusion\"]\nassert R[\"el_meniscus\"][\"Medial Meniscus\"] > 0.5, R[\"el_meniscus\"][\"Medial Meniscus\"]\n# severity is ordered, which is all AUC reads\nassert R[\"eff_large\"][\"Effusion\"] > R[\"eff_trace\"][\"Effusion\"] > R[\"eff_silent\"][\"Effusion\"], (\n    R[\"eff_large\"][\"Effusion\"], R[\"eff_trace\"][\"Effusion\"], R[\"eff_silent\"][\"Effusion\"])\n# a whole-joint statement reaches every compartment\nassert min(R[\"global_oa\"][t] for t in (\"Medial OA\", \"Lateral OA\", \"PF OA\")) > 0.5\n# reactive subchondral oedema must not read as a bruise as strongly as explicit trauma\nassert R[\"degen_marrow\"][\"Contusion\"] < R[\"true_bruise\"][\"Contusion\"], (\n    R[\"degen_marrow\"][\"Contusion\"], R[\"true_bruise\"][\"Contusion\"])\n\nprint(\"extractor self-test passed: negation, group statements, heading merge, Turkish and\")\nprint(\"Greek morphology, severity ordering, global OA propagation, contusion decoy\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 3. Reading the acquisition\n\n`train_series.csv` describes each series with an anatomical plane and two binary flags,\n`Fluid_Sensitive` and `Fat_Suppression`. They name two physically independent properties.\n\n*Fluid sensitivity* is a property of the **contrast weighting**, set by repetition time\n$T_R$ and echo time $T_E$. Long $T_R$ suppresses $T_1$ contrast and long $T_E$ builds\n$T_2$ contrast, giving the familiar three regimes:\n\n$$\n\\text{weighting} \\;=\\;\n\\begin{cases}\nT_1 & T_R \\lesssim 800\\ \\text{ms}\\\\\nT_2 & T_R \\gtrsim 800\\ \\text{ms},\\ T_E \\gtrsim 60\\ \\text{ms}\\\\\n\\text{PD} & T_R \\gtrsim 800\\ \\text{ms},\\ T_E \\lesssim 60\\ \\text{ms}\n\\end{cases}\n$$\n\nFluid is bright on $T_2$ and intermediate on proton density; on $T_1$ it is dark. Gradient\necho breaks the rule — its $T_R$ is short by design — so it is kept separate rather than\ncalled $T_1$.\n\n*Fat suppression* is a **preparation** applied on top of any weighting: a chemically\nselective pulse, an inversion (STIR), or water excitation. It is what makes marrow oedema\nconspicuous, because without it the bright fat signal hides it.\n\nThe two are orthogonal, and both are recoverable from the header independently of the\nprovided columns: $T_R$ and $T_E$ give the weighting, and `SeriesDescription`,\n`SequenceName` and `ScanOptions` name the suppression. Two cautions when reading those\nstrings, both of which silently invert the answer if missed:\n\n- underscore is a word character, so a token test for `we` (water excitation) never fires\n  inside `t2_de3d_we_tra`. Separators must be normalised to spaces first.\n- `ScanOptions` must be matched as exact tokens. One vendor writes `SAT_GEMS` for\n  *spatial* saturation, so a substring test for `SAT` marks non-fat-suppressed series as\n  suppressed.\n\n### Which sequences to show the model\n\nA knee is read in three planes because the structures run in different directions:\ncruciate ligaments obliquely, best seen sagittally; collateral ligaments and the meniscal\nbody coronally; patellar cartilage and the retinacula axially. Crossing plane with the two\nacquisition axes gives the slots below, chosen so that each of the twelve findings has at\nleast one sequence that shows it well.\n\n| slot | plane | weighting | fat sat | what it carries |\n|---|---|---|---|---|\n| `SAG_FLUID_FS` | sagittal | PD / T2 | yes | meniscal tears, marrow oedema, effusion |\n| `COR_FLUID_FS` | coronal | PD / T2 | yes | collateral ligaments, meniscal body, oedema |\n| `AX_FLUID_FS` | axial | PD / T2 | yes | patellofemoral joint, synovium, effusion |\n| `SAG_FLUID_NOFS` | sagittal | PD / T2 | no | meniscal morphology at high contrast-to-noise |\n| `COR_T1` | coronal | T1 | no | marrow architecture, cartilage and bone outline |\n| `SAG_T1` | sagittal | T1 | no | anatomy, chronic change |\n\nA study rarely has all six; a per-slot presence mask carries the absences into the head,\nwhich §6 uses.\n","metadata":{}},{"cell_type":"code","source":"from __future__ import annotations\n\nimport os\n\nfor _v in (\"OMP_NUM_THREADS\", \"OPENBLAS_NUM_THREADS\", \"MKL_NUM_THREADS\"):\n    os.environ.setdefault(_v, \"4\")\n\nimport gc\nimport re\nimport time\nimport traceback\nfrom concurrent.futures import ThreadPoolExecutor\nfrom pathlib import Path\n\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\n# The label extractor is defined in the cells above when this runs as a notebook. As a\n# plain script it is imported from the package source, so the two paths share one\n# definition rather than keeping a copy each.\nif \"extract\" not in dir():\n    import sys\n    for _p in (\"src\", \"../src\", \"../../src\"):\n        if (Path(_p) / \"report_labeler.py\").is_file():\n            sys.path.insert(0, _p)\n            break\n    from report_labeler import extract  # noqa: F401\n\nT0 = time.time()\nSEED = 2026\nnp.random.seed(SEED)\ntorch.manual_seed(SEED)\n\nTARGETS = [\"ACL\", \"MCL\", \"Medial Meniscus\", \"Lateral Meniscus\", \"Medial OA\",\n           \"Lateral OA\", \"PF OA\", \"Effusion\", \"Synovitis\", \"Baker's\",\n           \"Contusion\", \"Fracture\"]\n\n\nIMG = 224                  # encoder input; a multiple of the patch size 14\nCROP_MM = 160.0            # physical extent of the centre crop; a knee FOV is 140-180\nGROUP = 3                  # slices per encoder input, stacked as the three channels\nN_GROUP_MAX = 3            # groups per slot, before the memory budget is applied\nCACHE_BUDGET_GB = 12.0     # ceiling for the training cache\nHDR_THREADS = 16\nPIX_THREADS = 12\n\nEPOCHS = 12\nBATCH_STUDIES = 8          # a study is a bag of up to N_SLOT slot images\nLR_BACKBONE = 8e-6         # the encoder is being adapted, not retrained\nLR_HEAD = 1e-3\nWEIGHT_DECAY = 0.02\nUNFREEZE_LAST = 6          # trainable transformer blocks, counted from the output end\nEVAL_BATCH = 12\nTIME_BUDGET = 8.0 * 3600\n\n# Six slots: three planes crossed with the acquisition axes. The fat-suppressed\n# fluid-sensitive series exist for nearly every study; the T1 and the non-suppressed\n# fluid-sensitive series are scarcer, which is what the presence mask is for.\nSLOTS_RECOVERED = [\n    (\"SAG_FLUID_FS\", \"Sagittal\", True, True),\n    (\"COR_FLUID_FS\", \"Coronal\", True, True),\n    (\"AX_FLUID_FS\", \"Axial\", True, True),\n    (\"SAG_FLUID_NOFS\", \"Sagittal\", True, False),\n    (\"COR_T1\", \"Coronal\", False, False),\n    (\"SAG_T1\", \"Sagittal\", False, False),\n]\n\n# The alternative: plane x the provided flag, ignoring the recovered weighting. Kept as\n# a switch so the choice of slot definition can be varied while everything else is held\n# fixed. Under this scheme a `Struct` slot mixes T1 series with non-fat-suppressed PD/T2\n# series, which carry very different tissue contrast.\nSLOTS_PUBLIC = [\n    (\"SAG_FLUID\", \"Sagittal\", None, True),\n    (\"COR_FLUID\", \"Coronal\", None, True),\n    (\"AX_FLUID\", \"Axial\", None, True),\n    (\"SAG_STRUCT\", \"Sagittal\", None, False),\n    (\"COR_STRUCT\", \"Coronal\", None, False),\n    (\"AX_STRUCT\", \"Axial\", None, False),\n]\n\nSLOT_SCHEME = os.environ.get(\"SLOT_SCHEME\", \"recovered\")\nSLOTS = SLOTS_PUBLIC if SLOT_SCHEME == \"public\" else SLOTS_RECOVERED\nN_SLOT = len(SLOTS)\n\nFATSAT_OPTS = {\"FS\", \"FATSAT\", \"FAT_SAT\", \"FSAT\"}\n_SEP = re.compile(r\"[_\\-.]\")\n_FATSAT_RX = re.compile(r\"\\bfs\\b|fatsat|fat sat|\\bstir\\b|\\bspair\\b|\\bspir\\b|\\bwe\\b|\"\n                        r\"water excit|\\btirm\\b|\\bsting\\b|\\bfatsup\\b\")\n_T1_RX = re.compile(r\"\\bt1\\b|\\bt1w\\b\")\n_T2_RX = re.compile(r\"\\bt2\\b|\\bt2w\\b\")\n_PD_RX = re.compile(r\"\\bpd\\b|\\bpdw\\b|proton|\\bdp\\b|dens\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- v4 configuration ----------------------------------------------------------------\n# Only the knobs this version adds or changes. Everything above keeps its original value,\n# so setting all of these to their \"off\" state reproduces the previous pipeline exactly.\n\nimport math\nfrom copy import deepcopy\n\nimport matplotlib as mpl\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nN_FOLDS       = 4          # report-hash groups; folds are trained until the budget runs out\nMAX_FOLDS     = \"auto\"     # \"auto\" fits as many as the time budget allows, or an int\nFOLD_TIME_PAD = 1.15       # safety factor on the measured per-fold time\n\nUSE_VFLIP     = False      # was True; a knee is not vertically symmetric\nAUG_AFFINE    = True\nAUG_ROT_DEG   = 8.0\nAUG_SCALE     = 0.08\nAUG_SHIFT     = 0.05\nAUG_INTENSITY = 0.10\n\nEMA_DECAY     = 0.997      # 0 disables the moving average\nRANK_LOSS_W   = 0.05       # 0 disables the pairwise term\nRANK_POS, RANK_NEG = 0.60, 0.40      # graded-target cutoffs that define a usable pair\n\nLAT_FALLBACK        = \"auto\"   # \"auto\" | \"on\" | \"off\": infer side from x when the tag is absent\nLAT_MIN_AGREEMENT   = 0.85     # required agreement with the tag before \"auto\" enables it\nLAT_MIN_OFFSET_MM   = 5.0      # |x| below this is not a usable side cue\n\nSELECTION     = \"worse_of_two\"   # \"worse_of_two\" | \"derived\"\nGOLD_WEIGHT   = 3.0              # weight of an annotated study relative to a derived one\n\n# ------------------------------------------------------------------ icefire theme\nICEFIRE  = sns.color_palette(\"icefire\", 12)\nICE      = ICEFIRE[2]\nFIRE     = ICEFIRE[9]\nMIDLINE  = \"#39414d\"\nCMAP     = \"icefire\"\nCMAP_SEQ = sns.blend_palette([\"#eef1f6\", ICE, FIRE], as_cmap=True)\n\nsns.set_theme(style=\"white\", palette=ICEFIRE, font_scale=0.95)\nmpl.rcParams.update({\n    \"figure.dpi\": 120, \"savefig.dpi\": 120,\n    \"axes.spines.top\": False, \"axes.spines.right\": False,\n    \"axes.edgecolor\": \"#c8cdd4\", \"axes.labelcolor\": MIDLINE,\n    \"axes.titlesize\": 12, \"axes.titleweight\": \"bold\", \"axes.titlepad\": 10,\n    \"text.color\": MIDLINE, \"xtick.color\": MIDLINE, \"ytick.color\": MIDLINE,\n    \"grid.color\": \"#e8eaee\", \"axes.grid\": True, \"axes.grid.axis\": \"y\",\n    \"legend.frameon\": False, \"figure.facecolor\": \"white\",\n})\n\n\ndef grad(n, reverse=False):\n    \"\"\"n colours along the icefire ramp, skipping its near-black centre.\"\"\"\n    cm = plt.get_cmap(CMAP)\n    cold = np.linspace(0.02, 0.28, n - n // 2)\n    warm = np.linspace(0.78, 0.99, n // 2)\n    cols = [cm(p) for p in np.concatenate([cold, warm])][:n]\n    return cols[::-1] if reverse else cols\n\n\n# `log` is defined in the next cell, so this one reports with print().\nprint(f\"v4: folds={N_FOLDS} vflip={USE_VFLIP} ema={EMA_DECAY} rank_w={RANK_LOSS_W} \"\n      f\"lat_fallback={LAT_FALLBACK}\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def log(msg):\n    print(f\"[{time.time() - T0:7.1f}s] {msg}\", flush=True)\n\n\ndef find_root():\n    for c in [Path(\"/kaggle/input/competitions/rsna-knee-abnormality-detection\"),\n              Path(\"/kaggle/input/rsna-knee-abnormality-detection\"),\n              Path(\"data\"), Path(\".\")]:\n        if (c / \"test.csv\").is_file() and (c / \"test_series\").is_dir():\n            return c\n    # last resort: two-level scan, because the mount is nested one deeper than usual\n    for depth1 in sorted(p for p in Path(\"/kaggle/input\").iterdir() if p.is_dir()):\n        for cand in [depth1] + sorted(p for p in depth1.iterdir() if p.is_dir()):\n            if (cand / \"test.csv\").is_file():\n                return cand\n    raise FileNotFoundError(\"competition mount not found\")\n\n\ndef find_dinov2(variant=\"small\"):\n    \"\"\"Locate a mounted DINOv2 checkpoint directory by variant name.\"\"\"\n    base = Path(\"/kaggle/input\")\n    if not base.is_dir():\n        return None\n    hits = []\n    for root, dirs, files in os.walk(base):\n        dirs[:] = [d for d in dirs if d not in (\"train_series\", \"test_series\")]\n        if \"config.json\" in files and \"dinov2\" in root.lower():\n            hits.append(Path(root))\n    for h in hits:\n        if variant in str(h).lower():\n            return h\n    return hits[0] if hits else None\n\n\nROOT = find_root()\nlog(f\"input root: {ROOT}\")\n\n\ndef plan_cache(n_study):\n    \"\"\"Choose how many slices per slot the memory budget allows.\n\n    The cache is n_study x n_slot x slices x IMG^2 bytes. Coverage is the cheap axis -\n    linear - and resolution the expensive one, so when the budget binds it is the slice\n    count that gives way rather than the pixel grid. Deciding once, from the training\n    corpus size, keeps train and test caches on the same group layout.\n    \"\"\"\n    per_slice = n_study * N_SLOT * IMG * IMG\n    afford = int(CACHE_BUDGET_GB * 1024 ** 3 // max(per_slice, 1))\n    groups = max(1, min(N_GROUP_MAX, afford // GROUP))\n    if groups < N_GROUP_MAX:\n        log(f\"cache budget {CACHE_BUDGET_GB:.0f} GB allows {groups} group(s) of {GROUP}, \"\n            f\"not {N_GROUP_MAX}\")\n    return groups\n\n\nN_GROUP = plan_cache(len(pd.read_csv(ROOT / \"train.csv\")))\nCACHE_SLICES = GROUP * N_GROUP\nlog(f\"cache layout: {N_GROUP} groups x {GROUP} slices = {CACHE_SLICES} per slot\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Measuring the extractor before trusting it\n\nThe two gauges described above, run on the corpus. Neither can be skipped: the silence rate\nsays whether a rule *fires*, the agreement AUC says whether a fired rule is *right*, and\nonly the first has enough sample to decide anything.\n\nThe Hanley–McNeil interval is drawn on the agreement chart because it is the whole argument\nfor not tuning on that number. If the bar for a target is shorter than its interval, any\nchange that appears to improve it is indistinguishable from resampling noise.","metadata":{}},{"cell_type":"code","source":"train_df_diag = pd.read_csv(ROOT / \"train.csv\")\n_t = time.time()\nlab_diag = pd.DataFrame([extract(r) for r in train_df_diag[\"Report\"].fillna(\"\")])\nlog(f\"extracted {len(lab_diag)} reports in {time.time() - _t:.1f}s\")\n\nsilence = pd.Series(\n    {t: float(((lab_diag[t + \"__npos\"] == 0) & (lab_diag[t + \"__nneg\"] == 0)).mean())\n     for t in TARGETS})\n\ngold_diag = train_df_diag.set_index(\"StudyInstanceUID\")[TARGETS]\ngold_diag = gold_diag[gold_diag.notna().all(axis=1)]\n\n\ndef hanley_mcneil_se(auc, n_pos, n_neg):\n    \"\"\"Standard error of an AUC estimate; the reason the annotated subset cannot tune.\"\"\"\n    if not np.isfinite(auc) or n_pos < 1 or n_neg < 1:\n        return np.nan\n    q1 = auc / (2 - auc)\n    q2 = 2 * auc ** 2 / (1 + auc)\n    var = (auc * (1 - auc) + (n_pos - 1) * (q1 - auc ** 2)\n           + (n_neg - 1) * (q2 - auc ** 2)) / (n_pos * n_neg)\n    return float(np.sqrt(max(var, 0.0)))\n\n\nfrom sklearn.metrics import roc_auc_score\n\nlab_diag.index = train_df_diag[\"StudyInstanceUID\"].values\nrows = []\nfor t in TARGETS:\n    y = gold_diag[t].astype(int).values\n    p = lab_diag.loc[gold_diag.index, t].values\n    auc = roc_auc_score(y, p) if len(set(y)) > 1 else np.nan\n    se = hanley_mcneil_se(auc, int(y.sum()), int((1 - y).sum()))\n    rows.append({\"target\": t, \"agreement_auc\": auc, \"se\": se,\n                 \"n_pos\": int(y.sum()), \"n_neg\": int((1 - y).sum()),\n                 \"silence_rate\": silence[t]})\ndiag = pd.DataFrame(rows)\n\nfig, axes = plt.subplots(1, 2, figsize=(15, 4.4))\norder = diag.sort_values(\"agreement_auc\")\nax = axes[0]\nax.barh(order.target, order.agreement_auc, color=grad(12), edgecolor=\"white\", height=.7)\nax.errorbar(order.agreement_auc, np.arange(12), xerr=1.96 * order.se, fmt=\"none\",\n            ecolor=MIDLINE, elinewidth=1.2, capsize=3)\nax.axvline(.5, color=MIDLINE, ls=\":\", lw=1)\nax.set_xlim(0, 1.05); ax.set_xlabel(\"AUC of the extracted score against the annotation\")\nax.set_title(f\"Agreement, with 95% Hanley-McNeil intervals (n={len(gold_diag)})\")\nax.grid(axis=\"x\"); ax.grid(axis=\"y\", visible=False)\n\nax = axes[1]\norder2 = diag.sort_values(\"silence_rate\", ascending=False)\nax.barh(order2.target, order2.silence_rate * 100, color=grad(12, reverse=True),\n        edgecolor=\"white\", height=.7)\nax.set_xlabel(\"% of reports where no rule fired at all\")\nax.set_title(\"Silence rate — the gauge with enough sample to steer by\")\nax.grid(axis=\"x\"); ax.grid(axis=\"y\", visible=False)\nplt.tight_layout(); plt.show()\n\ndisplay(diag.sort_values(\"silence_rate\", ascending=False).round(3))\nprint(f\"mean agreement AUC {diag.agreement_auc.mean():.3f}, \"\n      f\"mean 95% half-width +-{1.96 * diag.se.mean():.3f}\")\nprint(\"A target whose interval is wider than the gap you are chasing cannot arbitrate a \"\n      \"lexicon change. Use the silence column for that.\")\n\ndel train_df_diag, lab_diag, gold_diag\ngc.collect()","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"HDR_TAGS = [\"SeriesDescription\", \"SequenceName\", \"ScanOptions\", \"ScanningSequence\",\n            \"RepetitionTime\", \"EchoTime\", \"Laterality\", \"ImageLaterality\",\n            \"ImagePositionPatient\", \"PixelSpacing\", \"Rows\",\n            \"Columns\", \"RescaleSlope\", \"RescaleIntercept\"]\n\n\ndef probe(item):\n    split, study, series, path = item\n    row = {\"split\": split, \"StudyInstanceUID\": study, \"SeriesInstanceUID\": series,\n           \"dir\": path}\n    try:\n        files = sorted(e.name for e in os.scandir(path) if e.name.endswith(\".dcm\"))\n        row[\"files\"] = files\n        row[\"n_slices\"] = len(files)\n        if not files:\n            return row\n        ds = pydicom.dcmread(os.path.join(path, files[len(files) // 2]),\n                             stop_before_pixels=True, force=True)\n        for t in HDR_TAGS:\n            v = getattr(ds, t, None)\n            if v is None:\n                row[t] = None\n            elif isinstance(v, (list, tuple)) or type(v).__name__ == \"MultiValue\":\n                row[t] = \"|\".join(str(x) for x in v)\n            else:\n                row[t] = str(v)\n    except Exception as exc:\n        row[\"err\"] = str(exc)[:120]\n    return row\n\n\ndef walk(split):\n    base = ROOT / split\n    items = []\n    if not base.is_dir():\n        return pd.DataFrame()\n    for study in os.scandir(base):\n        if study.is_dir():\n            for series in os.scandir(study.path):\n                if series.is_dir():\n                    items.append((split, study.name, series.name, series.path))\n    with ThreadPoolExecutor(max_workers=HDR_THREADS) as pool:\n        rows = list(pool.map(probe, items))\n    return pd.DataFrame(rows)\n\n\ndef annotate(df):\n    \"\"\"Recover fat suppression and pulse-sequence weighting from the header.\"\"\"\n    desc = (df[\"SeriesDescription\"].fillna(\"\") + \" \" + df[\"SequenceName\"].fillna(\"\"))\n    desc = desc.str.lower().str.replace(_SEP, \" \", regex=True)\n\n    opts = df[\"ScanOptions\"].fillna(\"\").str.upper().str.split(\"|\")\n    # GE writes SAT_GEMS for spatial saturation, so ScanOptions must be matched as\n    # exact tokens; a substring test on \"SAT\" fires on non-fat-sat series.\n    opts_fs = opts.apply(lambda ts: any(t.strip() in FATSAT_OPTS for t in ts))\n    df[\"fatsat\"] = desc.str.contains(_FATSAT_RX) | opts_fs\n\n    tr = pd.to_numeric(df[\"RepetitionTime\"], errors=\"coerce\")\n    te = pd.to_numeric(df[\"EchoTime\"], errors=\"coerce\")\n    gre = df[\"ScanningSequence\"].fillna(\"\").str.upper().str.contains(\"GR\")\n    t1, t2, pdw = desc.str.contains(_T1_RX), desc.str.contains(_T2_RX), desc.str.contains(_PD_RX)\n\n    df[\"weight\"] = np.where(t1 & ~t2 & ~pdw, \"T1\",\n                     np.where(t2 & ~pdw, \"T2\",\n                       np.where(pdw, \"PD\",\n                         np.where(gre, \"GRE\",\n                           np.where(tr < 800, \"T1\",\n                             np.where(te > 60, \"T2\",\n                               np.where(tr >= 800, \"PD\", \"UNK\")))))))\n    df[\"fluid\"] = np.isin(df[\"weight\"], [\"PD\", \"T2\"])\n    df[\"px\"] = pd.to_numeric(\n        df[\"PixelSpacing\"].fillna(\"\").str.split(\"|\").str[0].replace(\"\", np.nan),\n        errors=\"coerce\")\n    return df\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- laterality resolution -----------------------------------------------------------\n# The tag is authoritative where it exists. Where it does not, the patient x-coordinate\n# can stand in, but only if it agrees with the tag on the studies that have both: a wrong\n# flip is worse than no flip, and knee coils often place the joint near isocentre, where\n# the sign carries no information at all.\n\ndef _tag_side(g):\n    v = [str(x).strip().upper() for x in g[\"Laterality\"].dropna()]\n    if \"ImageLaterality\" in g.columns:\n        v += [str(x).strip().upper() for x in g[\"ImageLaterality\"].dropna()]\n    v = [x[0] for x in v if x and x[0] in (\"L\", \"R\")]\n    return v[0] if v else None\n\n\ndef _position_side(g):\n    xs = []\n    for s in g.get(\"ImagePositionPatient\", pd.Series(dtype=object)).dropna():\n        try:\n            xs.append(float(str(s).split(\"|\")[0]))\n        except Exception:\n            pass\n    if not xs:\n        return None\n    x = float(np.median(xs))\n    if abs(x) < LAT_MIN_OFFSET_MM:\n        return None                      # centred in the coil: the sign means nothing\n    return \"R\" if x < 0 else \"L\"         # LPS: the right knee sits at negative x\n\n\ndef laterality_maps(h):\n    \"\"\"Return (side_by_study, diagnostics). Uses the fallback only if it earns it.\"\"\"\n    tag, pos = {}, {}\n    for st, g in h.groupby(\"StudyInstanceUID\"):\n        tag[st] = _tag_side(g)\n        pos[st] = _position_side(g)\n\n    both = [st for st in tag if tag[st] and pos[st]]\n    agree = float(np.mean([tag[st] == pos[st] for st in both])) if both else np.nan\n    have_tag = float(np.mean([v is not None for v in tag.values()]))\n\n    if LAT_FALLBACK == \"on\":\n        use = True\n    elif LAT_FALLBACK == \"off\":\n        use = False\n    else:\n        use = bool(both) and np.isfinite(agree) and agree >= LAT_MIN_AGREEMENT\n\n    side = {st: (tag[st] or (pos[st] if use else None)) for st in tag}\n    covered = float(np.mean([v is not None for v in side.values()]))\n    info = {\"tag_coverage\": have_tag, \"agreement\": agree, \"n_compared\": len(both),\n            \"fallback_used\": use, \"final_coverage\": covered}\n    log(f\"laterality: tag on {have_tag:.1%} of studies, x-sign agrees with it on \"\n        f\"{agree:.1%} of {len(both)} comparable studies, fallback \"\n        f\"{'enabled' if use else 'disabled'} -> {covered:.1%} normalised\")\n    return side, info","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def pick_slots(series_df, plane_map):\n    \"\"\"One series per slot per study.\n\n    Ties are broken toward the stack with the most slices: a thicker stack samples the\n    joint more densely, and the three-slice sampler below benefits from the margin.\n    \"\"\"\n    series_df = series_df.copy()\n    series_df[\"plane\"] = series_df[\"SeriesInstanceUID\"].map(plane_map)\n    out = {}\n    for study, g in series_df.groupby(\"StudyInstanceUID\"):\n        chosen = {}\n        for name, plane, fluid, fs in SLOTS:\n            sel = (g[\"plane\"] == plane) & (g[\"fatsat\"] == fs)\n            # fluid=None means \"do not condition on weighting\" - the public scheme,\n            # where the single provided flag stands in for both axes at once.\n            if fluid is not None:\n                sel &= (g[\"fluid\"] == fluid)\n            cand = g[sel]\n            if len(cand) == 0 and fluid is False:\n                # T1 slots are the scarcest; fall back to any non-fat-sat, non-fluid\n                # series in the plane before giving up on the slot entirely.\n                cand = g[(g[\"plane\"] == plane) & (~g[\"fatsat\"])]\n            if len(cand):\n                chosen[name] = cand.sort_values(\"n_slices\", ascending=False).iloc[0]\n        out[study] = chosen\n    return out\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Sampling at a fixed physical scale\n\nA DICOM slice of $N \\times N$ pixels with spacing $s$ mm/pixel covers $N s$ millimetres of\nanatomy. Both $N$ and $s$ vary widely across this corpus, so a fixed-pixel resize —\n\"make everything $224 \\times 224$\" — hands the encoder images whose physical scale differs\nby a factor of about three. A meniscus then occupies a different number of pixels in\ndifferent studies for no anatomical reason, and the network has to spend capacity undoing\na nuisance transform it was handed for free.\n\nCrop to a constant physical extent $L$ before resizing. Taking the centre\n\n$$n \\;=\\; \\Big\\lfloor \\frac{L}{s} \\Big\\rceil \\quad\\text{pixels}$$\n\nand resampling that to $P \\times P$ leaves an effective scale of\n\n$$s_{\\text{eff}} \\;=\\; \\frac{L}{P}\\ \\ \\text{mm/pixel}$$\n\nwhich no longer depends on the acquisition. With $L = 160$ mm — a knee fits comfortably\ninside that — and $P = 224$, every study reaches the encoder at $0.71$ mm/pixel.\n\n**Intensity needs the same treatment for the same reason.** MR has no absolute scale: the\nsame tissue takes a different number on a different sequence, coil or day, so there is no\nHounsfield-unit equivalent to anchor to. Normalising each series to its own 1st and 99th\npercentile — over the whole volume, not per slice, so slices keep their relative contrast\n— removes an offset that would otherwise track the acquisition site. Percentiles rather\nthan min and max, because one bright vessel would otherwise compress everything else.\n\nSlices are taken from the central 60% of each stack. The outer slices of a knee series\nmostly lie outside the joint, and sampling them spends read bandwidth on soft tissue.\n","metadata":{}},{"cell_type":"code","source":"def read_slot(rec, n_slice=None, out_size=None):\n    \"\"\"`n_slice` physically spread slices from one series, at `out_size` pixels.\n\n    Returns float32 [n_slice, out, out] normalised per-series to its 1st-99th\n    percentile. Percentiles rather than min/max because MR intensity has no absolute\n    scale and a single bright vessel would otherwise compress the whole dynamic range.\n\n    Reading is the expensive half of this pipeline, so the caller reads once at the\n    largest configuration it needs and derives the smaller ones from the returned buffer\n    rather than re-reading.\n    \"\"\"\n    # was `N_SLICE`, which is never defined: a latent NameError on the default path.\n    n_slice = CACHE_SLICES if n_slice is None else n_slice\n    out_size = IMG if out_size is None else out_size\n    files, d, px = rec[\"files\"], rec[\"dir\"], rec[\"px\"]\n    n = len(files)\n    if n == 0:\n        return None\n    # Spread the samples over the central 60% of the stack: the outer slices of a knee\n    # series are mostly soft tissue outside the joint.\n    lo, hi = int(0.20 * (n - 1)), int(0.80 * (n - 1))\n    idx = np.unique(np.linspace(lo, hi, n_slice).astype(int)) if hi > lo else np.array([n // 2])\n    while len(idx) < n_slice:\n        idx = np.append(idx, idx[-1])\n\n    planes = []\n    for i in idx[:n_slice]:\n        try:\n            ds = pydicom.dcmread(os.path.join(d, files[int(i)]), force=True)\n            a = ds.pixel_array.astype(np.float32)\n            sl = float(getattr(ds, \"RescaleSlope\", 1) or 1)\n            ic = float(getattr(ds, \"RescaleIntercept\", 0) or 0)\n            a = a * sl + ic\n        except Exception:\n            a = np.zeros((out_size, out_size), dtype=np.float32)\n        planes.append(a)\n\n    shp = planes[0].shape\n    planes = [p if p.shape == shp else np.zeros(shp, np.float32) for p in planes]\n    vol = np.stack(planes)\n\n    # constant physical extent, then resize: PixelSpacing varies 3.4x across the corpus\n    if px and np.isfinite(px) and px > 0:\n        want = int(round(CROP_MM / px))\n        h, w = shp\n        if 16 < want < min(h, w):\n            cy, cx = h // 2, w // 2\n            half = want // 2\n            vol = vol[:, max(0, cy - half):cy + half, max(0, cx - half):cx + half]\n\n    lo_v, hi_v = np.percentile(vol, [1, 99])\n    vol = np.clip((vol - lo_v) / max(hi_v - lo_v, 1e-6), 0, 1)\n\n    t = torch.from_numpy(np.ascontiguousarray(vol)).unsqueeze(0)\n    t = F.interpolate(t, size=(out_size, out_size), mode=\"bilinear\", align_corners=False)\n    # uint8, not float32. These buffers queue up between the reader threads and the\n    # encoder, and at this size a float32 slot-series is several megabytes. Intensity is\n    # already normalised into [0, 1] here, so eight bits cost nothing that a bilinear\n    # resize has not already cost, and the queue is a quarter the size.\n    return (t.squeeze(0) * 255).round().clamp(0, 255).to(torch.uint8)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Normalising left and right\n\nFour of the twelve targets — the two menisci and the medial and lateral tibiofemoral\ncompartments — are medial/lateral pairs. Medial and lateral are defined relative to the\nbody's midline, so which side of the *image* they fall on depends on which knee was\nscanned. Unless that is normalised, those four labels are being asked to learn from an\naxis the model cannot observe.\n\nThe correction is not the same in every plane, because the mirror acts on a different\nimage axis:\n\n- **Coronal and axial.** The medial-lateral direction lies in the image plane, so a left\n  knee is the horizontal mirror of a right knee. Flipping the last axis maps one onto the\n  other.\n- **Sagittal.** The medial-lateral direction is the *slice* axis; each individual slice is\n  unchanged by mirroring. What differs is the order in which the stack traverses the\n  joint, so the slice order is reversed rather than the pixels flipped.\n\n`Laterality` is present in the header for some studies and absent for others. Where it is\nabsent the volume is left alone: a wrong flip is worse than no flip, and the presence mask\nlets the head learn how much to trust each slot.\n","metadata":{}},{"cell_type":"code","source":"def normalise_laterality(img, plane, lat):\n    \"\"\"Map every knee onto a left-knee convention.\n\n    Coronal and axial views mirror under a horizontal flip. Sagittal stacks are not\n    mirror images of each other - the slice order runs medial-to-lateral in opposite\n    directions - so the channel order is reversed instead.\n    \"\"\"\n    if lat != \"R\":\n        return img\n    if plane in (\"Coronal\", \"Axial\"):\n        return torch.flip(img, dims=[-1])\n    return torch.flip(img, dims=[0])\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 5. Reading once, training many times\n\nThe cost of this pipeline is dominated by reading, not by arithmetic. A study holds\nseveral series and each series holds tens of slices, so a study is on the order of a\nhundred and fifty files, and the corpus is hundreds of thousands of decodes.\n\nThat is affordable once. It is not affordable once per epoch, and fine-tuning needs the\nsame pixels every epoch. So the slot images are decoded a single time into memory and\nheld as `uint8`:\n\n$$\\text{bytes} \\;=\\; N_{\\text{study}} \\times N_{\\text{slot}} \\times S \\times P^{2}$$\n\nwith $S$ slices kept per slot at $P$ pixels. The exponent on $P$ is what makes this a\nreal constraint rather than a detail — the cache grows with the *square* of resolution\nand only linearly with slices, so coverage is the cheap axis and resolution the expensive\none. Eight bits cost nothing that the intensity normalisation of §4 has not already cost.\n\nThe cached slices are read as groups of three, each group forming one three-channel\nencoder input. Training draws one group per study per step, which doubles as augmentation\nalong the stack; inference averages the logits over all groups.\n\nTwo implementation consequences follow, both about the queue between the reading threads\nand the consumer rather than about either alone:\n\n- buffers crossing that queue are `uint8` for the same reason the cache is.\n- reads are issued in bounded chunks. Submitting every job at once lets the readers run\n  arbitrarily far ahead and the completed results accumulate without limit.\n\nA valid submission file is written before any of this begins and overwritten once real\npredictions exist, so the run always leaves a scoreable file behind.\n","metadata":{}},{"cell_type":"code","source":"def build_cache(slot_map, plane_map, lat_map, tag):\n    \"\"\"Decode every (study, slot) once into an in-memory uint8 array.\n\n    Fine-tuning revisits the same pixels every epoch. Reading them from the mount each\n    time would make the epoch count a function of I/O rather than of learning, so they\n    are decoded once and held as bytes: intensity has already been normalised into\n    [0, 1], and eight bits cost nothing a bilinear resize has not already cost.\n\n    CACHE_SLICES positions are kept per slot, which the training loop reads as N_GROUP\n    groups of GROUP consecutive channels.\n    \"\"\"\n    studies = sorted(slot_map)\n    sidx = {s: i for i, s in enumerate(studies)}\n    cache = np.zeros((len(studies), N_SLOT, CACHE_SLICES, IMG, IMG), np.uint8)\n    mask = np.zeros((len(studies), N_SLOT), np.float32)\n    log(f\"{tag}: cache {cache.shape} = {cache.nbytes / 1024 ** 3:.1f} GB\")\n\n    jobs = [(st, k, plane, slot_map[st][name])\n            for st in studies\n            for k, (name, plane, _, _) in enumerate(SLOTS)\n            if name in slot_map[st]]\n    log(f\"{tag}: decoding {len(jobs)} slot-series\")\n\n    CHUNK = 512\n    done = 0\n    with ThreadPoolExecutor(max_workers=PIX_THREADS) as pool:\n        for c0 in range(0, len(jobs), CHUNK):\n            block = jobs[c0:c0 + CHUNK]\n            for (st, k, plane, _), img in zip(\n                    block, pool.map(lambda j: read_slot(j[3], CACHE_SLICES, IMG), block)):\n                done += 1\n                if img is None:\n                    continue\n                cache[sidx[st], k] = normalise_laterality(img, plane,\n                                                          lat_map.get(st)).numpy()\n                mask[sidx[st], k] = 1.0\n            if done % 4096 < CHUNK:\n                log(f\"  {tag} {done}/{len(jobs)}\")\n            if time.time() - T0 > TIME_BUDGET:\n                log(f\"  {tag}: time budget reached during decode\")\n                break\n    gc.collect()\n    return studies, cache, mask\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Aggregating slots into twelve decisions\n\nA study arrives as up to six slot embeddings $x_s \\in \\mathbb{R}^{d}$ with a presence\nmask $m_s \\in \\{0,1\\}$. Pooling them identically would discard the reason the protocol\nhas three planes at all: each finding is read on particular sequences, and a mean over\nslots dilutes the one that carries the evidence with five that do not.\n\nProject each slot, add a learned slot identity, give every diagnosis $o$ its own query\n$q_o \\in \\mathbb{R}^{H}$, and let it attend over the slots with absent ones masked out of\nthe softmax:\n\n$$h_s \\;=\\; \\phi(x_s) + e_s, \\qquad\n\\alpha_{o,s} \\;=\\; \\frac{\\exp\\!\\big(\\langle h_s, q_o\\rangle / \\sqrt{H}\\big)\\, m_s}\n{\\sum_{s'} \\exp\\!\\big(\\langle h_{s'}, q_o\\rangle / \\sqrt{H}\\big)\\, m_{s'}},$$\n\n$$c_o \\;=\\; \\sum_s \\alpha_{o,s}\\, h_s, \\qquad\n\\ell_o \\;=\\; \\langle c_o, w_o \\rangle + b_o .$$\n\nThe masked softmax renormalises over whatever the study actually contains, so a missing\naxial series shifts a diagnosis's attention onto the sequences that are present instead\nof feeding it a zero vector.\n\n**The head is deliberately this small.** Richer aggregations are available — attention\nover every slice group rather than every slot, or a maximum instead of a mean — and they\nare worse. The reason is structural: the label is attached to the *study*, so nothing in\nthe supervision says which part of a study carries the finding. Extra attention\nparameters have no signal to learn that from, and spend their capacity on noise instead.\nWhere the supervision is coarse, the aggregation should be too.\n","metadata":{}},{"cell_type":"code","source":"class SlotHead(nn.Module):\n    \"\"\"Per-diagnosis attention over the slot embeddings of one study.\n\n    Each finding is read on particular sequences - cruciates sagittally, collateral\n    ligaments and the meniscal body coronally, patellar cartilage axially - so pooling\n    the slots identically would dilute the one that carries the evidence with the rest.\n\n    The aggregation is deliberately this simple. Richer alternatives were measured\n    against it and lost: with a study-level label there is no signal telling the model\n    which part of a study matters, so extra attention parameters have nothing to learn\n    from and spend their capacity fitting noise.\n    \"\"\"\n\n    def __init__(self, dim, n_slot, n_out, hidden=256, p=0.2):\n        super().__init__()\n        self.proj = nn.Sequential(nn.LayerNorm(dim), nn.Linear(dim, hidden), nn.GELU())\n        self.slot_emb = nn.Parameter(torch.randn(n_slot, hidden) * 0.02)\n        self.query = nn.Parameter(torch.randn(n_out, hidden) * 0.02)\n        self.drop = nn.Dropout(p)\n        self.out = nn.Linear(hidden, n_out)\n        self.hidden = hidden\n\n    def forward(self, x, mask):\n        h = self.proj(x) + self.slot_emb\n        att = torch.einsum(\"bsh,oh->bos\", h, self.query) / self.hidden ** 0.5\n        att = att.masked_fill(mask.unsqueeze(1) < 0.5, -1e4).softmax(-1)\n        ctx = self.drop(torch.einsum(\"bos,bsh->boh\", att, h))\n        return (ctx * self.out.weight.unsqueeze(0)).sum(-1) + self.out.bias\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class Model(nn.Module):\n    \"\"\"Encoder plus head, trained end to end.\n\n    A study arrives as a bag of slot images. The bag is flattened for the encoder and\n    folded back before the head, so the encoder never sees the study structure and the\n    head never sees pixels.\n    \"\"\"\n\n    def __init__(self, backbone, dim):\n        super().__init__()\n        self.backbone = backbone\n        self.head = SlotHead(dim, N_SLOT, len(TARGETS))\n        self.register_buffer(\"mean\", torch.tensor([0.485, 0.456, 0.406]).view(1, 3, 1, 1))\n        self.register_buffer(\"std\", torch.tensor([0.229, 0.224, 0.225]).view(1, 3, 1, 1))\n\n    def forward(self, imgs, mask):\n        B, S = imgs.shape[:2]\n        x = imgs.reshape(B * S, *imgs.shape[2:]).float().div_(255.0)\n        x = (x - self.mean) / self.std\n        out = self.backbone(pixel_values=x).last_hidden_state\n        feat = torch.cat([out[:, 0], out[:, 1:].mean(1)], dim=1).reshape(B, S, -1)\n        return self.head(feat, mask)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### Why the encoder is trained rather than frozen\n\nA frozen self-supervised encoder is the cheap option and it is where this pipeline\nstarted. It stops paying, and the way it stops is informative: resolution, encoder size,\nslice coverage and slot aggregation were each varied while everything else was held\nfixed, and none of them moved the score beyond what the validation noise allows.\n\nThat pattern is diagnostic. Those four axes change how much the model *looks*, how\nclosely it looks, and how it summarises what it saw — but none of them changes the\nvocabulary it looks with. When every one of them saturates at the same place, the binding\nconstraint is the representation itself. And there is an obvious reason to expect it\nhere: the encoder learned its features from natural images, where nothing resembles the\nsignal a torn meniscus makes on a proton-density sequence.\n\nSo the encoder is adapted, with two restraints.\n\n**Only the last blocks move.** The early blocks of a vision transformer are generic edge\nand texture filters; the late blocks are where semantics live. There is not enough\nsupervision here to improve the early ones and quite enough to damage them, so they stay\nfrozen and the last few are trained.\n\n**The encoder learns far more slowly than the head.** The head is random at\ninitialisation and has everything to learn; the encoder starts from a good solution and\nneeds only to be moved off it. A single learning rate would either leave the head\nuntrained or destroy the encoder in the first few hundred steps, so the two parameter\ngroups get rates two orders of magnitude apart.\n\nTargets remain the report-derived labels of §2, weighted by the extractor's\nper-study confidence, so a study whose report never mentions a finding pulls weakly on\nthat output rather than asserting a negative.\n","metadata":{}},{"cell_type":"code","source":"def build_model():\n    \"\"\"Load the encoder and open the last few blocks for training.\n\n    Only the last UNFREEZE_LAST blocks and the final norm are trainable. The early\n    blocks of a self-supervised transformer are generic edge and texture filters; there\n    is not enough supervision here to improve them and quite enough to damage them.\n    \"\"\"\n    from transformers import AutoModel\n    p = find_dinov2(\"small\")\n    if p is None:\n        raise FileNotFoundError(\"DINOv2 weights not attached\")\n    bb = AutoModel.from_pretrained(str(p))\n    n_layer = len(bb.encoder.layer)\n    for prm in bb.parameters():\n        prm.requires_grad = False\n    for blk in bb.encoder.layer[max(0, n_layer - UNFREEZE_LAST):]:\n        for prm in blk.parameters():\n            prm.requires_grad = True\n    for prm in bb.layernorm.parameters():\n        prm.requires_grad = True\n    dim = bb.config.hidden_size * 2\n    trainable = sum(p.numel() for p in bb.parameters() if p.requires_grad)\n    log(f\"backbone: {n_layer} blocks, last {UNFREEZE_LAST} trainable \"\n        f\"({trainable / 1e6:.1f}M params), feature dim {dim}\")\n    return Model(bb, dim)\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def take_group(cache_rows, g):\n    \"\"\"Slice GROUP consecutive channels out of the cached slices.\"\"\"\n    return cache_rows[:, :, g * GROUP:(g + 1) * GROUP]\n\n\n# `augment` is defined further down, after the reasoning for dropping the flip.\n\n\n@torch.no_grad()\ndef predict(model, cache, mask, idx, dev):\n    \"\"\"Average the logits over the groups of each slot.\n\n    Training sees one group at a time, which acts as augmentation along the stack;\n    inference averages over all of them, which is the same aggregation the frozen\n    pipeline used and the one that measured best.\n    \"\"\"\n    model.eval()\n    out = []\n    for b in range(0, len(idx), EVAL_BATCH):\n        sel = idx[b:b + EVAL_BATCH]\n        rows = torch.from_numpy(cache[sel]).to(dev)\n        m = torch.from_numpy(mask[sel]).to(dev)\n        acc = None\n        for g in range(N_GROUP):\n            with torch.autocast(\"cuda\", enabled=dev.type == \"cuda\"):\n                z = model(take_group(rows, g), m).float()\n            acc = z if acc is None else acc + z\n        out.append(torch.sigmoid(acc / N_GROUP).cpu().numpy())\n    return np.concatenate(out) if out else np.zeros((0, len(TARGETS)), np.float32)\n\n\ndef macro_auc(y, p):\n    from sklearn.metrics import roc_auc_score\n    return float(np.nanmean([roc_auc_score(y[:, j], p[:, j])\n                             if len(set(y[:, j])) > 1 else np.nan\n                             for j in range(y.shape[1])]))\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### The training-side changes\n\n**Augmentation that preserves anatomy.** The previous version flipped vertically. A knee is\nnot symmetric about that axis: the femur is above and the tibia below, and `Medial OA`,\n`Lateral OA` and `PF OA` are defined by which bone surface carries the cartilage loss.\nTeaching the encoder to ignore up from down removes information three targets depend on.\nRotation, scale and translation, all small, leave the anatomy intact while still preventing\nmemorisation of the exact framing. Horizontal flip stays forbidden for the reason given in\n§5 — it would undo the laterality normalisation.\n\n**An exponential moving average of the weights.** Both validation references are noisy, and\npicking a single epoch by them is close to picking a neighbour at random. An average over\nthe trajectory is more stable than any point on it, and it costs one extra copy of a model\nthat is mostly frozen anyway.\n\n**A pairwise ranking term.** The score is an AUC, and AUC counts correctly ordered\npositive–negative pairs. Cross-entropy is a proxy for that, and a reasonable one, but the\nquantity itself can be optimised directly: inside each batch, for each target, take the\nstudies whose graded target is confidently positive and confidently negative and push their\nlogits apart. Batches are small, so this fires for the common targets and stays silent for\nthe rare ones, which is why its weight is small and it supplements the cross-entropy rather\nthan replacing it.","metadata":{}},{"cell_type":"code","source":"def augment(imgs):\n    \"\"\"Small affine and intensity jitter over a whole bag. No flips of either kind.\n\n    Horizontal flip would undo the laterality normalisation of section 5. Vertical flip\n    would exchange the femoral and tibial surfaces, which is the axis the compartmental\n    osteoarthritis targets are defined on.\n    \"\"\"\n    B, S, C, H, W = imgs.shape\n    x = imgs.float()\n\n    if AUG_AFFINE:\n        ang = math.radians((torch.rand(1).item() - 0.5) * 2 * AUG_ROT_DEG)\n        sc = 1.0 + (torch.rand(1).item() - 0.5) * 2 * AUG_SCALE\n        tx = (torch.rand(1).item() - 0.5) * 2 * AUG_SHIFT\n        ty = (torch.rand(1).item() - 0.5) * 2 * AUG_SHIFT\n        cos, sin = math.cos(ang) / sc, math.sin(ang) / sc\n        theta = torch.tensor([[cos, -sin, tx], [sin, cos, ty]],\n                             dtype=torch.float32, device=imgs.device)\n        theta = theta.unsqueeze(0).repeat(B * S, 1, 1)\n        flat = x.reshape(B * S, C, H, W)\n        grid = F.affine_grid(theta, flat.shape, align_corners=False)\n        x = F.grid_sample(flat, grid, mode=\"bilinear\", padding_mode=\"zeros\",\n                          align_corners=False).reshape(B, S, C, H, W)\n\n    scale = 1.0 + (torch.rand(1, device=imgs.device) - 0.5) * 2 * AUG_INTENSITY\n    x = (x * scale).clamp(0, 255)\n\n    if USE_VFLIP and torch.rand(1).item() < 0.5:\n        x = torch.flip(x, dims=[-2])\n    return x.round().to(imgs.dtype)\n\n\nclass Ema:\n    \"\"\"Exponential moving average of the trainable weights.\n\n    Buffers are copied rather than averaged: the normalisation constants are fixed, and\n    averaging them would only accumulate floating point error.\n    \"\"\"\n\n    def __init__(self, model, decay):\n        self.decay = decay\n        self.step = 0\n        self.model = deepcopy(model).eval()\n        for p in self.model.parameters():\n            p.requires_grad_(False)\n\n    @torch.no_grad()\n    def update(self, model):\n        if self.decay <= 0:\n            return\n        # Warm-up: without it the average stays anchored to the random initialisation for\n        # the first few hundred steps, which is most of a short fold.\n        self.step += 1\n        decay = min(self.decay, (1.0 + self.step) / (10.0 + self.step))\n        src = model.state_dict()\n        for k, v in self.model.state_dict().items():\n            s = src[k]\n            if v.dtype.is_floating_point:\n                v.mul_(decay).add_(s.detach(), alpha=1.0 - decay)\n            else:\n                v.copy_(s)\n\n    def target(self, model):\n        \"\"\"The weights to evaluate and to keep: the average, or the model if disabled.\"\"\"\n        return self.model if self.decay > 0 else model\n\n\ndef rank_loss(logits, y, w):\n    \"\"\"Pairwise ranking term on confidently graded pairs inside the batch.\n\n    AUC counts ordered positive-negative pairs, so this optimises the metric directly.\n    A target with no usable pair in the batch contributes nothing rather than a constant.\n    \"\"\"\n    parts = []\n    usable = w > 0\n    for j in range(logits.shape[1]):\n        pos = logits[(y[:, j] > RANK_POS) & usable[:, j], j]\n        neg = logits[(y[:, j] < RANK_NEG) & usable[:, j], j]\n        if len(pos) and len(neg):\n            parts.append(F.softplus(-(pos[:, None] - neg[None, :])).mean())\n    return torch.stack(parts).mean() if parts else logits.new_tensor(0.0)\n\n\n# --- self-tests ----------------------------------------------------------------------\n_t = torch.randint(0, 255, (2, 3, GROUP, 32, 32), dtype=torch.uint8)\n_a = augment(_t)\nassert _a.shape == _t.shape and _a.dtype == _t.dtype\n# an asymmetric marker must not end up mirrored in either axis\n_m = torch.zeros(1, 1, GROUP, 32, 32, dtype=torch.uint8)\n_m[..., :6, :] = 255                                  # bright band along the top\n_out = augment(_m).float()\nassert _out[..., :16, :].mean() > _out[..., 16:, :].mean(), \"vertical orientation was lost\"\n_m2 = torch.zeros(1, 1, GROUP, 32, 32, dtype=torch.uint8)\n_m2[..., :6] = 255                                    # bright band on the left\n_out2 = augment(_m2).float()\nassert _out2[..., :16].mean() > _out2[..., 16:].mean(), \"horizontal orientation was lost\"\n\n_lg = torch.tensor([[2.0], [-2.0]])\n_y = torch.tensor([[1.0], [0.0]])\n_w = torch.ones(2, 1)\nassert rank_loss(_lg, _y, _w) < rank_loss(-_lg, _y, _w), \"ranking loss has the wrong sign\"\nassert float(rank_loss(_lg, torch.tensor([[1.0], [1.0]]), _w)) == 0.0, \"no pair, no loss\"\nprint(\"augmentation preserves both anatomical axes; ranking loss is oriented correctly\")","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 7. Validating without fooling yourself\n\nTwo leaks are specific to this setup, and both inflate a validation number without\nimproving anything.\n\n**Shared reports.** Some reports are byte-identical across studies — a template read for\nan unremarkable knee. Every study in such a group receives the same derived target vector.\nSplit the group across folds and the model is scored on a target whose source it has\nalready been trained on. Folds are therefore assigned by a hash of the report text, which\nkeeps every duplicate group whole inside one fold.\n\n**Two references, two meanings.** Out-of-fold performance is reported twice: against the\nderived targets, which covers every training study and measures whether the imaging model\nlearned what the text says; and against the studies with per-condition annotations, which\nis far smaller and measures agreement with a reading of the *images*. The second is the\none that resembles the test set. The first is the one with enough positives per label to\ndistinguish a real difference from noise.\n\n### Choosing which epoch to keep\n\nHaving two references makes selection a decision rather than a lookup, and the obvious\nrule — take the best epoch on the larger reference — is wrong. The two measure different\nthings, so an epoch that gains on one while losing on the other has not been shown to be\nbetter; it has been shown to be different. Ranking by\n\n$$\\text{score}(e) \\;=\\; \\min\\big(\\mathrm{AUC}^{\\text{derived}}_e,\\ \\mathrm{AUC}^{\\text{annot}}_e\\big)$$\n\nrefuses that trade: an epoch is only preferred if its *weaker* evidence is stronger than\nevery other epoch's weaker evidence. The checkpoint that wins under this rule is the one\nrestored before predicting, which also makes the training length a measured quantity\nrather than a hyperparameter guessed in advance.\n","metadata":{}},{"cell_type":"markdown","source":"### Two changes to how this is validated\n\n**The annotated reference was scoring studies the model had trained on.** The annotated\nsubset is small enough that all of it was previously handed to the epoch-selection rule,\nincluding the four fifths of it that sat in the training split. An epoch that had begun to\nmemorise those studies therefore looked better on the reference meant to detect exactly\nthat. Here each fold is scored only on the annotated studies held out of it, and the folds\ntogether produce one out-of-fold prediction per annotated study, so the final figure covers\nthe whole annotated subset without any of it having been trained on by the model that\npredicted it.\n\n**One fold became a cross-validation.** Training a single model on four fifths of the corpus\nwastes the remaining fifth and stakes everything on one trajectory. Each fold now\ncontributes a model, and the test predictions are combined as a rank mean per column —\nranks rather than probabilities, for the reason given in §1.\n\nHow many folds actually run is decided by the clock rather than in advance: the first fold\nis timed, and another is started only if the measured time plus a margin still fits inside\nthe budget. A submission is written after every fold, so a run that is cut short still\nleaves the ensemble of whatever finished.","metadata":{}},{"cell_type":"code","source":"def write_benchmark_submission():\n    \"\"\"Write the 0.5 benchmark file immediately.\n\n    A submission that never writes scores nothing at all. The try/except around main()\n    covers exceptions, but a kill for memory is a SIGKILL and never reaches it, so a valid\n    file exists from the first second and is overwritten only once predictions are ready.\n    \"\"\"\n    t = pd.read_csv(ROOT / \"test.csv\")\n    for c in TARGETS:\n        t[c] = 0.5\n    t.to_csv(\"submission.csv\", index=False)\n\n\ndef write_submission(rank_sum, n_models, st_te, test_df):\n    \"\"\"Rank-mean of the finished folds, aligned onto the test index.\"\"\"\n    P = rank_sum / max(n_models, 1)\n    sub = pd.DataFrame(P, columns=TARGETS)\n    sub.insert(0, \"StudyInstanceUID\", st_te)\n    sub = test_df[[\"StudyInstanceUID\"]].merge(sub, on=\"StudyInstanceUID\", how=\"left\")\n    sub[TARGETS] = sub[TARGETS].fillna(0.5)\n    sub.to_csv(\"submission.csv\", index=False)\n    return sub\n\n\ndef train_one_fold(fold, Ctr, Mtr, Y, W, tr, va, gi_va, gold_y_va, dev):\n    \"\"\"Train one fold and return (ema_state, history, val_predictions).\"\"\"\n    model = build_model().to(dev)\n    ema = Ema(model, EMA_DECAY)\n    opt = torch.optim.AdamW([\n        {\"params\": [p for p in model.backbone.parameters() if p.requires_grad],\n         \"lr\": LR_BACKBONE},\n        {\"params\": model.head.parameters(), \"lr\": LR_HEAD},\n    ], weight_decay=WEIGHT_DECAY)\n    steps = max(EPOCHS * (len(tr) // BATCH_STUDIES), 1)\n    sched = torch.optim.lr_scheduler.OneCycleLR(\n        opt, max_lr=[LR_BACKBONE, LR_HEAD], total_steps=steps, pct_start=0.15)\n    scaler = torch.amp.GradScaler(\"cuda\", enabled=dev.type == \"cuda\")\n\n    yv = (Y[va] > 0.5).astype(int)\n    best, best_state, history = -1.0, None, []\n\n    for ep in range(EPOCHS):\n        model.train()\n        perm = np.random.permutation(tr)\n        tot, nstep = 0.0, 0\n        for b in range(0, len(perm) - BATCH_STUDIES + 1, BATCH_STUDIES):\n            sel = perm[b:b + BATCH_STUDIES]\n            rows = torch.from_numpy(Ctr[sel]).to(dev)\n            g = int(torch.randint(N_GROUP, (1,)).item())\n            imgs = augment(take_group(rows, g))\n            m = torch.from_numpy(Mtr[sel]).to(dev)\n            y = torch.from_numpy(Y[sel]).to(dev)\n            w = torch.from_numpy(W[sel]).to(dev)\n            with torch.autocast(\"cuda\", enabled=dev.type == \"cuda\"):\n                z = model(imgs, m)\n                loss = (F.binary_cross_entropy_with_logits(z, y, reduction=\"none\") * w).mean()\n                if RANK_LOSS_W > 0:\n                    loss = loss + RANK_LOSS_W * rank_loss(z.float(), y, w)\n            opt.zero_grad(set_to_none=True)\n            scaler.scale(loss).backward()\n            scaler.step(opt)\n            scaler.update()\n            sched.step()\n            ema.update(model)\n            tot += loss.item()\n            nstep += 1\n\n        eval_model = ema.target(model)\n        pv = predict(eval_model, Ctr, Mtr, va, dev)\n        d = macro_auc(yv, pv)\n\n        # The annotated reference for this fold is only the annotated studies held out of\n        # it. Scoring the ones it trained on is what made the previous rule optimistic.\n        g_auc = float(\"nan\")\n        if gold_y_va is not None and len(gi_va) >= 6:\n            g_auc = macro_auc(gold_y_va, predict(eval_model, Ctr, Mtr, gi_va, dev))\n\n        history.append({\"fold\": fold, \"epoch\": ep + 1, \"loss\": tot / max(nstep, 1),\n                        \"derived\": d, \"annot\": g_auc})\n        log(f\"fold {fold} epoch {ep + 1}/{EPOCHS}  loss {tot / max(nstep, 1):.4f}\"\n            f\"  derived {d:.4f}  annot(held out) {g_auc:.4f}\")\n\n        if SELECTION == \"derived\" or not np.isfinite(g_auc):\n            score = d\n        else:\n            score = min(d, g_auc)\n        if score > best:\n            best = score\n            best_state = {k: v.detach().cpu().clone()\n                          for k, v in eval_model.state_dict().items()}\n            log(f\"  best so far ({SELECTION} {score:.4f})\")\n        if time.time() - T0 > TIME_BUDGET:\n            log(\"time budget reached inside fold\")\n            break\n\n    if best_state is None:\n        best_state = {k: v.detach().cpu().clone()\n                      for k, v in ema.target(model).state_dict().items()}\n    del model, opt, sched, scaler\n    gc.collect()\n    if torch.cuda.is_available():\n        torch.cuda.empty_cache()\n    return best_state, history, best\n\n\ndef main():\n    write_benchmark_submission()\n    test_df = pd.read_csv(ROOT / \"test.csv\")\n    test_series = pd.read_csv(ROOT / \"test_series.csv\")\n    train_df = pd.read_csv(ROOT / \"train.csv\")\n    train_series = pd.read_csv(ROOT / \"train_series.csv\")\n    log(f\"train {train_df.shape} test {test_df.shape}\")\n\n    both = pd.concat([train_series, test_series])\n    plane_map = dict(zip(both[\"SeriesInstanceUID\"], both[\"Anatomical_Plane\"]))\n\n    log(\"header pass: test\")\n    hte = annotate(walk(\"test_series\"))\n    log(f\"  {len(hte)} test series\")\n    log(\"header pass: train\")\n    htr = annotate(walk(\"train_series\"))\n    log(f\"  {len(htr)} train series\")\n\n    lat_tr, lat_info_tr = laterality_maps(htr)\n    lat_te, _ = laterality_maps(hte)\n    globals()[\"LAT_INFO\"] = lat_info_tr\n\n    slots_te, slots_tr = pick_slots(hte, plane_map), pick_slots(htr, plane_map)\n    cov = pd.Series([len(v) for v in slots_tr.values()]).describe()\n    log(f\"train slots per study: mean {cov['mean']:.2f} min {cov['min']:.0f} \"\n        f\"max {cov['max']:.0f}\")\n    globals()[\"SLOT_COVERAGE\"] = pd.Series(\n        {name: float(np.mean([name in v for v in slots_tr.values()]))\n         for name, *_ in SLOTS})\n    globals()[\"WEIGHT_TABLE\"] = htr.groupby([\"weight\", \"fatsat\"]).size()\n\n    st_tr, Ctr, Mtr = build_cache(slots_tr, plane_map, lat_tr, \"train\")\n    st_te, Cte, Mte = build_cache(slots_te, plane_map, lat_te, \"test\")\n\n    # ---- targets ---------------------------------------------------------- #\n    t_lab = time.time()\n    lab = pd.DataFrame([extract(r) for r in train_df[\"Report\"].fillna(\"\")])\n    lab[\"StudyInstanceUID\"] = train_df[\"StudyInstanceUID\"].values\n    lab = lab.set_index(\"StudyInstanceUID\")\n    log(f\"derived labels for {len(lab)} studies in {time.time() - t_lab:.1f}s\")\n\n    gold = train_df.set_index(\"StudyInstanceUID\")[TARGETS]\n    gold = gold[gold.notna().all(axis=1)]\n\n    Y = np.zeros((len(st_tr), len(TARGETS)), np.float32)\n    W = np.zeros_like(Y)\n    for i, st in enumerate(st_tr):\n        if st in gold.index:\n            Y[i], W[i] = gold.loc[st].values, GOLD_WEIGHT\n        elif st in lab.index:\n            r = lab.loc[st]\n            Y[i] = r[TARGETS].values\n            W[i] = 0.25 + 0.75 * r[[t + \"__conf\" for t in TARGETS]].values\n    keep = np.where(W.sum(1) > 0)[0]\n    log(f\"supervised {len(keep)} of {len(st_tr)} studies (annotated {len(gold)})\")\n\n    # Folds are grouped on report text: some reports are byte-identical across studies and\n    # yield one target vector for all of them, so splitting such a group would score a\n    # model on a target whose source it had already trained on.\n    import hashlib\n    rep = train_df.set_index(\"StudyInstanceUID\")[\"Report\"].fillna(\"\")\n    grp = np.array([int(hashlib.md5(rep.get(s, s).encode()).hexdigest()[:8], 16) % N_FOLDS\n                    for s in st_tr])\n    gpos = {s: i for i, s in enumerate(st_tr)}\n    gi_all = np.array([gpos[s] for s in gold.index if s in gpos])\n    gold_y_all = (gold.loc[[st_tr[i] for i in gi_all]].values.astype(int)\n                  if len(gi_all) else None)\n\n    dev = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n    rank_sum = np.zeros((len(st_te), len(TARGETS)), np.float64)\n    n_models, sub = 0, None\n    oof_gold = np.full((len(st_tr), len(TARGETS)), np.nan, np.float32)\n    histories, fold_scores = [], []\n\n    max_folds = N_FOLDS if MAX_FOLDS == \"auto\" else int(MAX_FOLDS)\n    for fold in range(min(N_FOLDS, max_folds)):\n        va = np.array([i for i in keep if grp[i] == fold])\n        tr = np.array([i for i in keep if grp[i] != fold])\n        if len(va) == 0 or len(tr) < BATCH_STUDIES:\n            log(f\"fold {fold}: not enough studies, skipped\")\n            continue\n        gi_va = np.array([i for i in gi_all if grp[i] == fold])\n        gold_y_va = (gold.loc[[st_tr[i] for i in gi_va]].values.astype(int)\n                     if len(gi_va) else None)\n        log(f\"=== fold {fold}: train {len(tr)} / holdout {len(va)} \"\n            f\"(annotated held out: {len(gi_va)}) ===\")\n\n        t_fold = time.time()\n        state, history, best = train_one_fold(fold, Ctr, Mtr, Y, W, tr, va,\n                                              gi_va, gold_y_va, dev)\n        fold_time = time.time() - t_fold\n        histories.extend(history)\n        fold_scores.append(best)\n\n        model = build_model().to(dev)\n        model.load_state_dict(state)\n        if len(gi_va):\n            oof_gold[gi_va] = predict(model, Ctr, Mtr, gi_va, dev)\n        P = predict(model, Cte, Mte, np.arange(len(st_te)), dev)\n        rank_sum += pd.DataFrame(P).rank(pct=True).values\n        n_models += 1\n        del model, state\n        gc.collect()\n        if torch.cuda.is_available():\n            torch.cuda.empty_cache()\n\n        sub = write_submission(rank_sum, n_models, st_te, test_df)\n        log(f\"fold {fold} done in {fold_time / 60:.1f} min; submission has {n_models} model(s)\")\n\n        elapsed = time.time() - T0\n        if elapsed + FOLD_TIME_PAD * fold_time > TIME_BUDGET:\n            log(f\"stopping after {n_models} fold(s): another would need \"\n                f\"{FOLD_TIME_PAD * fold_time / 60:.0f} min and \"\n                f\"{(TIME_BUDGET - elapsed) / 60:.0f} min remain\")\n            break\n\n    globals()[\"HISTORY\"] = pd.DataFrame(histories)\n    if gold_y_all is not None and len(gi_all):\n        seen = np.isfinite(oof_gold[gi_all]).all(axis=1)\n        if seen.sum() >= 8:\n            oof_auc = macro_auc(gold_y_all[seen], oof_gold[gi_all][seen])\n            log(f\"out-of-fold macro AUC on {int(seen.sum())} annotated studies: {oof_auc:.4f}\")\n            globals()[\"OOF_GOLD\"] = (gold_y_all[seen], oof_gold[gi_all][seen])\n    log(f\"fold selection scores: {[round(s, 4) for s in fold_scores]}\")\n    if sub is None:\n        # No fold finished. The benchmark file written at the top is still on disk, so the\n        # run leaves a scoreable submission rather than nothing.\n        log(\"no fold completed; the 0.5 benchmark submission stands\")\n        return\n    log(f\"submission.csv {sub.shape}; models {n_models}\")\n    print(sub.head().to_string())","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"try:\n    main()\nexcept Exception:\n    traceback.print_exc()\n    # A submission that fails to write scores nothing at all, so fall back to the\n    # benchmark file rather than dying.\n    t = pd.read_csv(find_root() / \"test.csv\")\n    for c in TARGETS:\n        t[c] = 0.5\n    t.to_csv(\"submission.csv\", index=False)\n    print(\"wrote fallback submission.csv\")\nlog(\"done\")\n","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --- what the run actually did -------------------------------------------------------\n# Reads the globals main() left behind. Safe to run after a partial run; anything the run\n# did not reach is simply skipped.\n\nif \"HISTORY\" in globals() and len(HISTORY):\n    h = HISTORY\n    fig, axes = plt.subplots(1, 2, figsize=(15, 4.2))\n    for k, f in enumerate(sorted(h.fold.unique())):\n        sub_h = h[h.fold == f]\n        axes[0].plot(sub_h.epoch, sub_h.derived, \"o-\", color=grad(max(len(h.fold.unique()), 3))[k],\n                     label=f\"fold {f}\")\n        axes[1].plot(sub_h.epoch, sub_h.annot, \"o-\", color=grad(max(len(h.fold.unique()), 3))[k],\n                     label=f\"fold {f}\")\n    axes[0].set_title(\"Derived-target holdout AUC\"); axes[0].set_xlabel(\"epoch\")\n    axes[1].set_title(\"Annotated holdout AUC (held out of this fold)\"); axes[1].set_xlabel(\"epoch\")\n    for ax in axes:\n        ax.legend(fontsize=8)\n    plt.tight_layout(); plt.show()\n    print(\"The right panel is the honest one and it is noisy by construction: roughly a\")\n    print(\"dozen annotated studies per fold. It guards the selection; it does not drive it.\")\n\nif \"OOF_GOLD\" in globals():\n    y, p = OOF_GOLD\n    rows = []\n    for j, t in enumerate(TARGETS):\n        if len(set(y[:, j])) > 1:\n            auc = roc_auc_score(y[:, j], p[:, j])\n            se = hanley_mcneil_se(auc, int(y[:, j].sum()), int((1 - y[:, j]).sum()))\n        else:\n            auc, se = np.nan, np.nan\n        rows.append({\"target\": t, \"oof_auc\": auc, \"se\": se, \"n_pos\": int(y[:, j].sum())})\n    oof_table = pd.DataFrame(rows)\n    fig, ax = plt.subplots(figsize=(10, 4))\n    o = oof_table.sort_values(\"oof_auc\")\n    ax.barh(o.target, o.oof_auc, color=grad(12), edgecolor=\"white\", height=.7)\n    ax.errorbar(o.oof_auc, np.arange(len(o)), xerr=1.96 * o.se, fmt=\"none\",\n                ecolor=MIDLINE, elinewidth=1.2, capsize=3)\n    ax.axvline(.5, color=MIDLINE, ls=\":\", lw=1)\n    ax.set_xlim(0, 1.05)\n    ax.set_title(f\"Out-of-fold AUC against the annotations, macro {oof_table.oof_auc.mean():.4f}\")\n    ax.grid(axis=\"x\"); ax.grid(axis=\"y\", visible=False)\n    plt.tight_layout(); plt.show()\n    display(oof_table.round(3))\n\nif \"SLOT_COVERAGE\" in globals():\n    fig, ax = plt.subplots(figsize=(8, 3.2))\n    sc = SLOT_COVERAGE.sort_values(ascending=False)\n    ax.bar(sc.index, sc.values * 100, color=grad(len(sc)), edgecolor=\"white\", width=.6)\n    ax.set_ylim(0, 112); ax.set_ylabel(\"% of studies\"); ax.set_title(\"Slot availability\")\n    ax.tick_params(axis=\"x\", rotation=25)\n    for i, v in enumerate(sc.values * 100):\n        ax.text(i, v + 2.5, f\"{v:.0f}\", ha=\"center\", fontsize=9, color=MIDLINE)\n    plt.tight_layout(); plt.show()\n\nif \"WEIGHT_TABLE\" in globals():\n    print(\"series by recovered weighting and fat suppression:\")\n    display(WEIGHT_TABLE.unstack(fill_value=0))\n\nif \"LAT_INFO\" in globals():\n    print(\"laterality:\", {k: (round(v, 3) if isinstance(v, float) else v)\n                          for k, v in LAT_INFO.items()})","metadata":{},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Where the next points are\n\n**The silence chart is the work queue.** A target that is silent on a large share of reports\nis one whose training signal is mostly the 0.28 floor, and no amount of encoder capacity\nrecovers a label that was never extracted. Read twenty reports in the language with the\nhighest silence for the worst target and the missing vocabulary is usually three or four\nstems.\n\n**The annotator's threshold is a target-specific offset, not a global one.** Section 2 grades\nmentions because the reporting radiologist and the annotator disagree about what counts.\nThat disagreement is not the same size for every finding — a trace effusion is routinely\nreported and rarely annotated, whereas a fracture is either there or not. The per-target\nagreement chart shows where the mismatch is worst, and the severity cutoffs in `_severity`\nare the knob for it.\n\n**Resolution is the untested axis.** The physical-scale argument in §4 fixes mm per pixel at\n`CROP_MM / IMG`. Both terms are free: a tighter crop at the same grid, or the same crop on a\nfiner grid, each buy resolution where the meniscus is. The cache is quadratic in `IMG` and\nlinear in slices, so the trade is against slice coverage, and the budget planner already\nmakes that trade explicit.\n\n**A second encoder would ensemble well.** The folds currently differ only by their data\nsplit. Swapping the backbone for one fold — a larger DINOv2, or a supervised ImageNet\ntransformer — decorrelates the errors far more than a different split does, and the rank\nmean already in place is the right way to combine them.\n\n**What not to spend time on.** Calibration, thresholds, and class rebalancing: §1 rules all\nthree out. And no horizontal or vertical flip anywhere — the first undoes the laterality\nnormalisation, the second exchanges the femur with the tibia.","metadata":{}}]}