{"cells":[{"cell_type":"markdown","id":"baa9a4cc","metadata":{},"source":"# `train.csv` Has 4,407 Studies and 58 Labels\n\n**`df.dropna()` leaves you with 58 rows. The other 4,349 studies aren't missing data —\ntheir labels are in the report text.**\n\nThis notebook counts what's actually in `train.csv`. That's all it does. No model, no\nextraction method, no strategy — but the count changes what kind of problem this is, and\nit takes one `value_counts()` to see.\n\nFour things it establishes:\n\n1. **58 of 4,407 studies (1.3%) carry labels.** Not partially — those 58 have all twelve\n   filled, and the rest have all twelve empty. It's all-or-nothing per study.\n2. **Every study has a report**, including all 4,349 unlabelled ones. The diagnostic\n   information is present; it just isn't in the label columns.\n3. **The corpus is multi-script** — 541 reports (12.3%) contain Greek or Cyrillic — and the\n   file has more lines than rows because report text contains newlines, so a naive line\n   count overstates the dataset by 13×.\n4. **New in v2, and it partly corrects this notebook:** section 3 warned that a\n   Latin-script text pipeline fails *silently* on those 541 reports. That was an argument,\n   never a measurement. Section 4 measures it against a published LLM labelling that\n   encodes refusals explicitly — and **the refusals go the other way.** 25.9% of label\n   cells are refused on Latin-only reports against 21.9% on non-Latin ones, a gap no\n   shuffle out of 2,000 reached."},{"cell_type":"code","execution_count":1,"id":"222c3ed5","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:12.94903Z","iopub.status.busy":"2026-08-15T16:52:12.948909Z","iopub.status.idle":"2026-08-15T16:52:13.633953Z","shell.execute_reply":"2026-08-15T16:52:13.633674Z"}},"outputs":[],"source":"import csv, collections, re, pathlib\n\ndef _safe_roots():\n    \"\"\"Search roots for shipped data. Deliberately excludes /, home and the\n    system temp dirs: rglob on those crawls the whole filesystem, which turns a\n    missing-data case into a hang (or an OSError on macOS AppTranslocation).\"\"\"\n    import os\n    here = pathlib.Path.cwd().resolve()\n    bad = {pathlib.Path(\"/\"), pathlib.Path.home(), pathlib.Path(\"/tmp\"),\n           pathlib.Path(\"/private\"), pathlib.Path(\"/private/tmp\"), pathlib.Path(\"/var\")}\n    roots = [pathlib.Path(\"/kaggle/input\"), here, *list(here.parents)[:4]]\n    return [r for r in roots if r.exists() and r not in bad]\n\nimport matplotlib.pyplot as plt\nimport numpy as np\n\ndef find_train():\n    for root in _safe_roots():\n        for p in root.rglob(\"train.csv\"):\n            try: head = p.open(encoding=\"utf-8\", errors=\"ignore\").readline()\n            except Exception: continue\n            # Validate it's THIS competition's file -- several ship a train.csv.\n            if \"StudyInstanceUID\" in head and \"Report\" in head:\n                return p\n    return None\n\npath = find_train()\nif path is None:\n    raise FileNotFoundError(\"Attach the RSNA knee competition data.\")\nprint(f\"reading {path}\")\n\nwith path.open(encoding=\"utf-8\") as fh:\n    n_lines = sum(1 for _ in fh)\nrows = list(csv.DictReader(path.open(encoding=\"utf-8\")))\nLABELS = [c for c in rows[0] if c not in (\"StudyInstanceUID\", \"Report\")]\n\nprint(f\"\\nlines in file : {n_lines:,}\")\nprint(f\"parsed studies: {len(rows):,}\")\nprint(f\"ratio         : {n_lines/len(rows):.1f} lines per study  <- reports contain newlines\")\nprint(f\"\\nlabel columns ({len(LABELS)}): {', '.join(LABELS)}\")"},{"cell_type":"markdown","id":"d61481c5","metadata":{},"source":"## 1. How many studies actually carry labels\n\nThe label columns are empty strings, not zeros — which matters, because an empty string is\n\"unknown\", while a zero is a negative finding. Conflating them silently invents ~4,349\nnegative examples per class that nobody asserted."},{"cell_type":"code","execution_count":2,"id":"174afc56","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:13.635298Z","iopub.status.busy":"2026-08-15T16:52:13.635198Z","iopub.status.idle":"2026-08-15T16:52:13.643323Z","shell.execute_reply":"2026-08-15T16:52:13.643107Z"}},"outputs":[],"source":"def is_labelled(r):  return any(r[l] != \"\" for l in LABELS)\n\nlabelled   = [r for r in rows if is_labelled(r)]\nunlabelled = [r for r in rows if not is_labelled(r)]\n\nprint(f\"total studies       {len(rows):>6,}\")\nprint(f\"labelled            {len(labelled):>6,}   {100*len(labelled)/len(rows):>5.1f}%\")\nprint(f\"unlabelled          {len(unlabelled):>6,}   {100*len(unlabelled)/len(rows):>5.1f}%\")\n\n# All-or-nothing, or partially filled?\nfilled = collections.Counter(sum(1 for l in LABELS if r[l] != \"\") for r in labelled)\nprint(f\"\\nlabel columns filled per labelled study: {dict(sorted(filled.items()))}\")\nprint(\"-> labelling is all-or-nothing per study, never partial.\")\n\nhave_report = sum(1 for r in rows if r[\"Report\"].strip())\nprint(f\"\\nstudies with a report: {have_report:,} ({100*have_report/len(rows):.1f}%)\")\nprint(\"Every unlabelled study still has its report text.\")"},{"cell_type":"markdown","id":"6dd02382","metadata":{},"source":"## 2. What the 58 labelled studies look like\n\nWorth knowing before treating them as a validation set: they are **not** a random sample.\nPrevalence in them is far higher than any knee-MRI population would show, which is what\nyou'd expect from a deliberately enriched annotation set."},{"cell_type":"code","execution_count":3,"id":"ec4f0ff1","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:13.644524Z","iopub.status.busy":"2026-08-15T16:52:13.644441Z","iopub.status.idle":"2026-08-15T16:52:13.718605Z","shell.execute_reply":"2026-08-15T16:52:13.718423Z"}},"outputs":[],"source":"prev = {}\nfor l in LABELS:\n    vals = [r[l] for r in labelled if r[l] != \"\"]\n    pos = sum(1 for v in vals if v == \"1\")\n    prev[l] = 100*pos/len(vals)\n\norder = sorted(prev, key=prev.get, reverse=True)\nprint(f\"{'finding':<20} {'positive':>9} {'of':>4} {'rate':>7}\")\nprint(\"-\"*44)\nfor l in order:\n    vals = [r[l] for r in labelled if r[l] != \"\"]\n    pos = sum(1 for v in vals if v == \"1\")\n    print(f\"{l:<20} {pos:>9} {len(vals):>4} {prev[l]:>6.1f}%\")\n\nfig, ax = plt.subplots(figsize=(9.5, 4.6))\nax.barh(order[::-1], [prev[l] for l in order[::-1]], color=\"#0074D9\")\nax.set_xlabel(\"% positive among the 58 labelled studies\")\nax.set_title(\"Prevalence in the labelled subset (not population prevalence)\", fontweight=\"bold\")\nax.spines[[\"top\",\"right\"]].set_visible(False)\nplt.tight_layout(); plt.show()\n\nn_pos = [sum(1 for l in LABELS if r[l]==\"1\") for r in labelled]\nprint(f\"\\nfindings per labelled study: mean {np.mean(n_pos):.1f}, max {max(n_pos)}\")\nprint(\"Multi-label and heavily co-occurring -- these are not mutually exclusive classes.\")"},{"cell_type":"markdown","id":"dcebbb41","metadata":{},"source":"## 3. The corpus is multi-script, and my first pass at this was wrong\n\n**Correcting v1 of this notebook.** It classified the reports with a keyword heuristic\nthat tested only English and Spanish terms, and bucketed everything else as \"ambiguous\".\nThat heuristic was structurally incapable of finding a third language, so the ambiguous\nbucket was not genuinely undecidable — it was hiding real languages, and it was large.\nThe notebook `RSNA Knee: read the report, then the knee` by prvsiyan mentions Greek\nreports, which is what tipped me off; I had not considered Greek at all.\n\nRedone by Unicode script instead of keywords. A report counts as containing a script if\nit has three or more consecutive letters of it, and a report can contain more than one:"},{"cell_type":"code","execution_count":4,"id":"b6a548f0","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:13.719823Z","iopub.status.busy":"2026-08-15T16:52:13.719736Z","iopub.status.idle":"2026-08-15T16:52:13.796691Z","shell.execute_reply":"2026-08-15T16:52:13.796455Z"}},"outputs":[],"source":"# \"3+ consecutive letters of the script\" -- deliberately strict, so a stray symbol\n# (a lone mu, a degree sign) does not classify a report as Greek.\nRUN = {\n    \"Latin\":    re.compile(r\"[A-Za-zÀ-ɏ]{3,}\"),\n    \"Greek\":    re.compile(r\"[Ͱ-Ͽἀ-῿]{3,}\"),\n    \"Cyrillic\": re.compile(r\"[Ѐ-ӿ]{3,}\"),\n}\n\nscript_of = {name: [bool(rx.search(r[\"Report\"])) for r in rows] for name, rx in RUN.items()}\nnon_latin = [g or c for g, c in zip(script_of[\"Greek\"], script_of[\"Cyrillic\"])]\n\nfor name in (\"Latin\", \"Greek\", \"Cyrillic\"):\n    n = sum(script_of[name])\n    print(f\"  {name:<9} {n:>5,}  {100*n/len(rows):>5.1f}%\")\nprint(f\"  {'non-Latin':<9} {sum(non_latin):>5,}  {100*sum(non_latin)/len(rows):>5.1f}%   <- Greek or Cyrillic present\")\n\n# The Greek is prose, not stray symbols: check that every Greek-flagged report qualifies.\nassert all(RUN[\"Greek\"].search(r[\"Report\"]) for r, g in zip(rows, script_of[\"Greek\"]) if g)\nprint(\"\\nOne Greek report opening, for illustration:\")\ngreek_ex = next(r for r, g in zip(rows, script_of[\"Greek\"]) if g)\nprint(\" \", greek_ex[\"Report\"][:90].replace(\"\\n\", \" \"))\n\nlens = sorted(len(r[\"Report\"]) for r in rows)\nprint(f\"\\nreport length (chars): median {lens[len(lens)//2]:,}  p90 {lens[int(.9*len(lens))]:,}  max {lens[-1]:,}\")\nprint(\"\\nfirst report, truncated:\")\nprint(\" \", rows[0][\"Report\"][:220].replace(\"\\n\", \" \"))"},{"cell_type":"markdown","id":"7c4a6691","metadata":{},"source":"This is a script check, not language identification — Latin covers English, Spanish and\nanything else in that alphabet — so treat it as a lower bound on the diversity present.\n\nOne codepoint trap while you are here: **U+03BC GREEK SMALL LETTER MU and U+00B5 MICRO\nSIGN render identically and do not match each other.** A regex for one silently misses\nthe other.\n\n## 4. Does a text pipeline go quiet on those reports? Measured — and it does not\n\nSection 3 is a warning as it stands: a pipeline built on Latin-script terms does not fail\nloudly on the 541 non-Latin reports, it extracts nothing and returns a confident negative.\nSame shape as the empty-versus-zero trap in section 1 — a silent wrong answer rather than\nan error.\n\n**That was an argument, never a measurement, and it is now testable.** stevenleehans\npublished an LLM labelling of all 4,407 reports\n([dataset](https://www.kaggle.com/datasets/stevenleehans/rsna-knee-llm-report-labels),\n[discussion](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection/discussion/733932))\nwhich encodes *\"the report does not address this\"* as an explicit **0.5** rather than\nguessing. That sentinel is a refusal counter. If a text pipeline were quietly failing on\nthe non-Latin reports, its refusals should pile up there.\n\nReproducing both source numbers first, as hard assertions — if either fails, the join is\nagainst the wrong file and nothing below means anything:"},{"cell_type":"code","execution_count":5,"id":"d7ed249d","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:13.797933Z","iopub.status.busy":"2026-08-15T16:52:13.797843Z","iopub.status.idle":"2026-08-15T16:52:13.934147Z","shell.execute_reply":"2026-08-15T16:52:13.933905Z"}},"outputs":[],"source":"def find_llm_labels():\n    # Exact filename: this dataset also ships llm_labels_v2.csv and\n    # llm_labels_v4_blend.csv, which are imputed and would give different numbers.\n    for root in _safe_roots():\n        for p in root.rglob(\"llm_labels_full.csv\"):\n            return p\n    return None\n\nlpath = find_llm_labels()\nif lpath is None:\n    raise FileNotFoundError(\n        \"Attach the dataset stevenleehans/rsna-knee-llm-report-labels (llm_labels_full.csv).\")\nprint(f\"reading {lpath}\")\n\nlrows = list(csv.DictReader(lpath.open(encoding=\"utf-8\")))\nLFIND = [c for c in lrows[0] if c != \"StudyInstanceUID\"]\nassert set(LFIND) == set(LABELS), \"label columns differ between the two files\"\n\nSENTINEL = 0.5\ncells = [[float(r[c]) for c in LFIND] for r in lrows]\nn_cells = sum(len(row) for row in cells)\nn_sent  = sum(1 for row in cells for v in row if v == SENTINEL)\n\nprint(f\"\\nCHECK 1 - refusal sentinel (published as 25.4% of all label cells)\")\nprint(f\"  cells at exactly 0.5: {n_sent:,} / {n_cells:,} = {100*n_sent/n_cells:.1f}%\")\nassert n_sent == 13438 and n_cells == 52884, \"sentinel count does not reproduce -- wrong file?\"\n\nprint(f\"\\nCHECK 2 - script counts from section 3 above\")\nprint(f\"  Greek {sum(script_of['Greek']):,} / Cyrillic {sum(script_of['Cyrillic']):,} \"\n      f\"/ non-Latin {sum(non_latin):,}\")\nassert sum(script_of[\"Greek\"]) == 321 and sum(script_of[\"Cyrillic\"]) == 220\n\nby_uid = {r[\"StudyInstanceUID\"]: r for r in lrows}\nmatched = [(r, by_uid[r[\"StudyInstanceUID\"]]) for r in rows if r[\"StudyInstanceUID\"] in by_uid]\nprint(f\"\\nCHECK 3 - join coverage\")\nprint(f\"  matched {len(matched):,} / {len(rows):,} studies ({100*len(matched)/len(rows):.1f}%)\")\nassert len(matched) == len(rows), \"incomplete join -- the groups below would not be comparable\""},{"cell_type":"markdown","id":"5b4b8f39","metadata":{},"source":"Now the cross-tab. Every study carries all twelve finding cells, so the *finding mix is\nidentical across the two groups by construction* — there is no composition confound to\nargue about and the two rates compare directly."},{"cell_type":"code","execution_count":6,"id":"d45fb619","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:13.935464Z","iopub.status.busy":"2026-08-15T16:52:13.935391Z","iopub.status.idle":"2026-08-15T16:52:13.942517Z","shell.execute_reply":"2026-08-15T16:52:13.94232Z"}},"outputs":[],"source":"groups = {\"Latin only\": [], \"non-Latin present\": []}\nfor (t, l), nl in zip(matched, non_latin):\n    groups[\"non-Latin present\" if nl else \"Latin only\"].append(l)\n\ndef refusal(rs):\n    vals = [float(r[c]) for r in rs for c in LFIND]\n    return sum(1 for v in vals if v == SENTINEL) / len(vals), len(vals)\n\nrates = {}\nfor g, rs in groups.items():\n    rate, n = refusal(rs)\n    rates[g] = rate\n    print(f\"  {g:<20} {len(rs):>5,} studies   refused {100*rate:>5.1f}% of {n:,} cells\")\n\ngap = rates[\"non-Latin present\"] - rates[\"Latin only\"]\nprint(f\"\\n  ratio (non-Latin / Latin-only): {rates['non-Latin present']/rates['Latin only']:.3f}x\")\nprint(f\"  absolute difference           : {100*gap:+.1f} pp\")\nprint(\"\\n  -> It refuses LESS on the non-Latin reports, not more.\")\n\nprint(\"\\nSplit by individual script (the groups overlap -- a report can contain both):\")\nfor name in (\"Greek\", \"Cyrillic\"):\n    rs = [l for (t, l), f in zip(matched, script_of[name]) if f]\n    print(f\"  {name:<9} {len(rs):>5,} studies   refused {100*refusal(rs)[0]:>5.1f}%\")"},{"cell_type":"markdown","id":"44740d5e","metadata":{},"source":"### Is that bigger than noise, and is it just report length?\n\nTwo checks. The permutation test shuffles the script label across all 4,407 studies and\nrecomputes the gap, which is the null \"script has nothing to do with refusal\". The length\ncheck is the mundane explanation — a longer report gives a labeller more to work with."},{"cell_type":"code","execution_count":7,"id":"ced243ef","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:13.943599Z","iopub.status.busy":"2026-08-15T16:52:13.943523Z","iopub.status.idle":"2026-08-15T16:52:14.206227Z","shell.execute_reply":"2026-08-15T16:52:14.205811Z"}},"outputs":[],"source":"rng = np.random.default_rng(0)\nflat = np.array([[float(l[c]) == SENTINEL for c in LFIND] for t, l in matched])\nmask = np.array(non_latin)\nobs  = flat[mask].mean() - flat[~mask].mean()\n\nidx  = np.arange(len(flat))\nnull = np.empty(2000)\nfor i in range(2000):\n    rng.shuffle(idx)\n    m = np.zeros(len(flat), dtype=bool); m[idx[:mask.sum()]] = True\n    null[i] = flat[m].mean() - flat[~m].mean()\n\nprint(f\"  observed difference {100*obs:+.2f} pp\")\nprint(f\"  null sd             {100*null.std():.2f} pp\")\nprint(f\"  shuffles reaching it: {int((np.abs(null) >= abs(obs)).sum())} / 2,000\")\n\nprint(\"\\nReport length by script group:\")\nfor g, sel in ((\"Latin only\", ~mask), (\"non-Latin present\", mask)):\n    ls = np.array([len(t[\"Report\"]) for t, l in matched])[sel]\n    print(f\"  {g:<20} median {int(np.median(ls)):>5,} chars   mean {ls.mean():>7,.0f}\")\nprint(\"\\n  -> The non-Latin reports are SHORTER, which predicts MORE refusals, not fewer.\")\nprint(\"     Length runs against this result rather than explaining it.\")"},{"cell_type":"markdown","id":"9333e18b","metadata":{},"source":"### Per finding — and the two columns that go the other way\n\nTen of the twelve findings are refused less often on non-Latin reports. Two are refused\n*more*, and they are already the two most-refused columns:"},{"cell_type":"code","execution_count":8,"id":"078aa715","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:14.207629Z","iopub.status.busy":"2026-08-15T16:52:14.207515Z","iopub.status.idle":"2026-08-15T16:52:14.266619Z","shell.execute_reply":"2026-08-15T16:52:14.266384Z"}},"outputs":[],"source":"per = []\nfor c in LFIND:\n    a = np.mean([float(l[c]) == SENTINEL for (t, l), f in zip(matched, non_latin) if not f])\n    b = np.mean([float(l[c]) == SENTINEL for (t, l), f in zip(matched, non_latin) if f])\n    per.append((c, a, b, b - a))\nper.sort(key=lambda r: r[3])\n\nfor c, a, b, d in per:\n    print(f\"  {c:<18} Latin {100*a:>5.1f}%   non-Latin {100*b:>5.1f}%   gap {100*d:>+5.1f} pp\")\n\nfig, ax = plt.subplots(figsize=(9.5, 5.0))\nnames = [c for c, a, b, d in per]\ngaps  = [100*d for c, a, b, d in per]\nax.barh(names, gaps, color=[\"#D62728\" if g > 0 else \"#0074D9\" for g in gaps])\nax.axvline(0, color=\"black\", lw=1)\nax.axvline(100*null.std()*2, color=\"grey\", ls=\"--\", lw=1)\nax.axvline(-100*null.std()*2, color=\"grey\", ls=\"--\", lw=1)\nax.set_xlabel(\"refusal rate on non-Latin reports minus Latin-only (percentage points)\")\nax.set_title(\"Blue = refused less on non-Latin reports. Red = refused more.\\n\"\n             \"Dashed lines: ±2 permutation SD of the overall gap\", fontweight=\"bold\")\nax.spines[[\"top\", \"right\"]].set_visible(False)\nplt.tight_layout(); plt.show()"},{"cell_type":"markdown","id":"6eff2d1e","metadata":{},"source":"**The caveat is the actionable half, so please don't take only the headline.**\n\nA lower refusal rate is **not** evidence of a correct read. A labeller that reads Greek\nfluently and one that hallucinates fluently produce the identical signature here: fewer\nrefusals. Refusal rate measures willingness to answer, not accuracy.\n\nAnd the accuracy question is not currently answerable, because the only ruler is the\nannotator-labelled studies:"},{"cell_type":"code","execution_count":9,"id":"f91041b7","metadata":{"execution":{"iopub.execute_input":"2026-08-15T16:52:14.267887Z","iopub.status.busy":"2026-08-15T16:52:14.267815Z","iopub.status.idle":"2026-08-15T16:52:14.271731Z","shell.execute_reply":"2026-08-15T16:52:14.271553Z"}},"outputs":[],"source":"gold_non_latin = sum(1 for r, nl in zip(rows, non_latin) if is_labelled(r) and nl)\nprint(f\"  gold-labelled studies       : {len(labelled):>3,}\")\nprint(f\"  of which non-Latin script   : {gold_non_latin:>3,}\")\nprint(\"\\n  -> That is the entire ruler available for checking accuracy on the 541\")\nprint(\"     non-Latin studies. It is not enough to settle it either way.\")"},{"cell_type":"markdown","id":"93dd54e4","metadata":{},"source":"So the claim section 4 actually supports, and no more than this: **the specific failure\nmode section 3 warns about — a pipeline going quiet on non-Latin reports — does not appear\nin this label set.** Whether the answers it gives on those 541 studies are *right* is a\nseparate question, and 6 gold studies cannot settle it.\n\nOne thing that may matter more than the headline, for anyone reading topic 733932: its\nsection 4 turns on how differently Synovitis and Baker's behave when a report is silent.\nThose are exactly the two columns whose silence rate shifts by script above, so that\nargument is being measured across a mixed population — worth checking whether it holds\nwithin script group."},{"cell_type":"markdown","id":"90f8e017","metadata":{},"source":"## What this changes\n\nWith 58 labelled studies across 12 correlated findings, there is not enough supervision to\ntrain a twelve-class imaging model directly, and not enough to validate one either — a\nheld-out split of 58 gives single-digit positives for the rarer findings.\n\nWhat there *is*: **4,407 free-text reports covering every study, including the 4,349\nwithout labels.** The diagnostic information exists for the whole training set; it's in\nprose rather than in columns.\n\nI'm deliberately not proposing how to use that — this notebook is a count, not a method.\nBut the count is worth having before choosing an approach, because the obvious first move\n(`dropna()`, then train) quietly discards 98.7% of the available signal.\n\nThree smaller traps in the same file:\n\n- **Empty is not zero.** Filling blanks with 0 invents thousands of unasserted negatives.\n- **58,556 lines, 4,407 studies.** Reports contain newlines, so `wc -l` overstates the\n  dataset 13×. Parse with a real CSV reader.\n- **U+03BC and U+00B5** both render as µ and do not match each other.\n\nAnd one thing section 4 changes about the report text specifically: if your pipeline drops\nor down-weights reports it cannot parse, you are discarding 541 studies (12.3%) that at\nleast one published labeller did *not* choke on. Whether it read them *correctly* is\nunsettled and, with 6 non-Latin studies among the 58 gold, currently unsettleable.\n\n---\n\n*Every number computed in-cell from `train.csv` and `llm_labels_full.csv`. The file lookup\nvalidates the header before using it, because several competitions in a shared workspace\nship a `train.csv`, and the three reproduction checks in section 4 are `assert`s rather\nthan prints so a wrong file stops the notebook instead of quietly changing the answer.*\n\n*v2 (2026-08-15) replaces the keyword language heuristic with a Unicode-script count and\nadds section 4. The refusal analysis was posted to topic 734055 the same day; credit to\nstevenleehans for publishing a label set with an explicit refusal code, which is the only\nreason any of section 4 is measurable, and to prvsiyan for the Greek observation.*\n\n*Corrections welcome — particularly on the script regexes, which are ranges rather than\nproper `unicodedata` script lookups.*"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.12.7"}},"nbformat":4,"nbformat_minor":5}