{"cells":[{"cell_type":"markdown","metadata":{},"source":"# RSNA step-3: a paired, same-split test of report-derived pseudo-labels\n\n**Question:** this competition scores macro ROC-AUC across 12 targets, but only **58 of 4,407**\ntrain studies carry real (gold) labels -- the rest have a radiology report and nothing else. Three\npublic Kaggle datasets (`pilkwang/rsna-knee-llm-labels`, `stevenleehans/rsna-knee-llm-report-labels`,\n`lixin73/rsna-knee-llm-report-labels-sol56`) each independently extracted per-target verdicts from\nthose reports with an LLM. **Does training on the studies where all three extractors agree actually\nhelp, or is it just more noise?**\n\n**Design, stated before running:** for each of 5 folds, train two models from the identical random\ninitialization and the identical fold split -- one on gold-labeled studies only (**baseline**), one\non gold plus 250 pseudo-labeled studies sampled from the >=8/12-label-unanimous pool (**treatment**)\n-- and evaluate both on the same held-out gold studies. The only thing that differs between the pair\nis whether pseudo-labeled studies were added to training. This is a stronger test than comparing two\nseparate runs: fold composition and initialization can't be confounds here, only the treatment can\nbe.\n\n**Grouping/leakage note (L3):** validation is 100% gold-only in every fold; pseudo-labeled studies\nare added to every fold's training set unconditionally and never validated on, so their fold tag\ncarries no leakage risk either way. The fold split itself is computed **fresh, in this kernel**, from\n`train.csv` (the standard, sanctioned `competition_sources` mount) -- patients are 1:1 with studies\nin this corpus (verified in the public notebook `wguesdon/rsna-knee-dinov2-at-meniscus-resolution`,\n4,407/4,407), so `StudyInstanceUID` is already the natural leakage-safe grouping unit and no external\nfold-assignment dataset is needed.\n\n**Masking, not filtering:** a pseudo-label is only used in the loss for a given (study, label) pair\nwhen all three independent sources agree; everywhere else that label is masked out for that study.\nStandard masked weak supervision.\n\n**Read this honestly before any number below:** validation is ~11-12 gold studies per fold, and this\nis a single N_PSEUDO=250 draw from the 2,570-study candidate pool. A result in either direction here\nis directional evidence about whether this technique is worth pursuing further at scale -- not a\nconfident, final claim either way. The title of the published version of this notebook reflects\nwhatever this run actually measured, decided after the run, not before it."},{"cell_type":"code","metadata":{},"source":"from __future__ import annotations\nimport gc\nimport time\nimport warnings\nfrom pathlib import Path\n\nimport cv2\nimport numpy as np\nimport pandas as pd\nimport pydicom\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.model_selection import KFold\n\nwarnings.filterwarnings('ignore')\ncv2.setNumThreads(1)\n\nT0 = time.time()\n\nCOMP = Path('/kaggle/input/competitions/rsna-knee-abnormality-detection')\nBACKBONE_PT = Path('/kaggle/input/datasets/marwanmath/resnet-50-radimagenet-marwan/ResNet50.pt')\nREPORT_PIL_CSV = Path('/kaggle/input/datasets/pilkwang/rsna-knee-llm-labels/report_labels_v2.csv')\nREPORT_STV_CSV = Path('/kaggle/input/datasets/stevenleehans/rsna-knee-llm-report-labels/llm_labels_v2.csv')\nREPORT_LIX_CSV = Path('/kaggle/input/datasets/lixin73/rsna-knee-llm-report-labels-sol56/labels_llm_gpt56sol.csv')\n# NOTE: no external fold-assignment dataset is mounted here on purpose -- see the markdown cell\n# above. Folds are computed inline from train.csv in the next cell.\n\nDEV = 'cuda' if torch.cuda.is_available() else 'cpu'\nassert DEV == 'cuda', 'GPU requested but not visible; refusing silent CPU fallback'  # L5/L6\nprint(f'device      : {DEV}')\nfor i in range(torch.cuda.device_count()):\n    cc = torch.cuda.get_device_capability(i)\n    print(f'  gpu{i}       : {torch.cuda.get_device_name(i)} sm_{cc[0]}{cc[1]}, '\n          f'{torch.cuda.get_device_properties(i).total_memory / 2 ** 30:.0f} GiB')\n\nassert COMP.exists(), f'competition data not mounted at {COMP}'\nassert BACKBONE_PT.is_file(), f'backbone checkpoint not found at {BACKBONE_PT}'\nassert REPORT_PIL_CSV.is_file(), f'pilkwang report-label file not found at {REPORT_PIL_CSV}'\nassert REPORT_STV_CSV.is_file(), f'stevenleehans report-label file not found at {REPORT_STV_CSV}'\nassert REPORT_LIX_CSV.is_file(), f'lixin73 report-label file not found at {REPORT_LIX_CSV}'\nprint('all paths resolved')\n\nTARGETS = ['ACL', 'MCL', 'Medial Meniscus', 'Lateral Meniscus', 'Medial OA', 'Lateral OA',\n           'PF OA', 'Effusion', 'Synovitis', \"Baker's\", 'Contusion', 'Fracture']\nCROP_MM = 130.0\nSIZE = 224                      # matches the RadImageNet encoder's proven input convention\nN_SLICE = 6                     # downscaled -- feature-extraction pass, not full-res training\nSLICE_BAND = (0.2, 0.8)\nSLOTS = [('Sagittal', 1), ('Sagittal', 0), ('Coronal', 1), ('Coronal', 0), ('Axial', 1), ('Axial', 0)]\nN_SLOT = len(SLOTS)\nSEED = 2026\nN_PSEUDO = 250                  # sampled from the >=8/12-unanimous bucket (2,570 candidates)\nMIN_LABELS_AGREED = 8           # per-study bar for inclusion in the sampling pool (not a per-label bar)\nN_FOLDS = 5\ntorch.manual_seed(SEED)\nnp.random.seed(SEED)","outputs":[],"execution_count":null},{"cell_type":"code","metadata":{},"source":"# --- DICOM -> image tensor extraction, reused verbatim in method from rsna-tonylica-repro\n# (already-proven, already-scored-0.920 code path) -- only SIZE/N_SLICE are downscaled for this\n# cheaper feature-extraction pass. Points at train_series/ instead of test_series/ since gold\n# labels (and the report-derived pseudo-labels) only exist for train-split studies.\n\nSERIES_ROOT = COMP / 'train_series'\nassert SERIES_ROOT.exists(), f'{SERIES_ROOT} missing'\n\n\ndef ordered_files(sdir, cap=64):\n    keyed = []\n    for f in sdir.glob('*.dcm'):\n        try:\n            ds = pydicom.dcmread(str(f), stop_before_pixels=True)\n            keyed.append((int(ds.InstanceNumber), str(f)))\n        except Exception:\n            continue\n        if len(keyed) >= cap * 4:\n            break\n    return [f for _, f in sorted(keyed)]\n\n\ndef series_side(path):\n    try:\n        return float(pydicom.dcmread(path, stop_before_pixels=True).ImagePositionPatient[0])\n    except Exception:\n        return 0.0\n\n\ndef read_crop(path):\n    try:\n        ds = pydicom.dcmread(path)\n        arr = ds.pixel_array.astype(np.float32)\n    except Exception:\n        return None\n    try:\n        ps = float(ds.PixelSpacing[0])\n    except Exception:\n        ps = CROP_MM / max(arr.shape)\n    half = int(round(CROP_MM / ps / 2))\n    cy, cx = (arr.shape[0] // 2, arr.shape[1] // 2)\n    y0, y1 = (max(0, cy - half), min(arr.shape[0], cy + half))\n    x0, x1 = (max(0, cx - half), min(arr.shape[1], cx + half))\n    crop = arr[y0:y1, x0:x1]\n    return None if crop.size == 0 else crop\n\n\ndef window(crop, lo, hi, flip):\n    c = np.clip((crop - lo) / max(hi - lo, 1e-06), 0, 1)\n    img = cv2.resize(c, (SIZE, SIZE), interpolation=cv2.INTER_AREA)\n    return img[:, ::-1].copy() if flip else img\n\n\ndef render(path, flip):\n    crop = read_crop(path)\n    if crop is None:\n        return None\n    lo, hi = np.percentile(crop[::4, ::4], [1, 99])\n    return window(crop, lo, hi, flip)\n\n\ndef build_study(study, recs):\n    \"\"\"Returns (N_SLOT, N_SLICE, SIZE, SIZE) uint8 tensor + (N_SLOT,) populated-slice mask.\"\"\"\n    out = np.zeros((N_SLOT, N_SLICE, SIZE, SIZE), np.uint8)\n    mask = np.zeros(N_SLOT, np.uint8)\n    rows = pd.DataFrame(recs)\n    if len(rows):\n        for s_i, (plane, fs) in enumerate(SLOTS):\n            sub = rows[(rows.Anatomical_Plane == plane) & (rows.Fat_Suppression == fs)]\n            if sub.empty:\n                continue\n            files = ordered_files(SERIES_ROOT / study / sub.iloc[0].SeriesInstanceUID)\n            if not files:\n                continue\n            flip = plane != 'Sagittal' and series_side(files[0]) < 0\n            lo, hi = SLICE_BAND\n            i0 = int(round(lo * (len(files) - 1)))\n            i1 = int(round(hi * (len(files) - 1)))\n            avail = list(range(i0, i1 + 1))\n            if len(avail) >= N_SLICE:\n                picks = [avail[int(round(t))] for t in np.linspace(0, len(avail) - 1, N_SLICE)]\n                off = 0\n            else:\n                picks, off = (avail, (N_SLICE - len(avail)) // 2)\n            for c, p in enumerate(picks):\n                img = render(files[p], flip)\n                if img is None:\n                    img = render(files[min(len(files) - 1, p + 1)], flip)\n                if img is not None:\n                    out[s_i, off + c] = (img * 255).astype(np.uint8)\n            mask[s_i] = len(picks)\n    return out, mask\n\n\nprint('DICOM extraction functions ready (reused from the proven 0.920 pipeline)')","outputs":[],"execution_count":null},{"cell_type":"code","metadata":{},"source":"# --- gold-labeled studies + a fresh, seeded, in-kernel grouped fold split.\n# No external fold-assignment dataset is read here -- see the markdown cell above for why.\n\ntrain = pd.read_csv(COMP / 'train.csv')\nser = pd.read_csv(COMP / 'train_series.csv')\nser = ser.loc[:, ~ser.columns.duplicated()]\n\ngold = train[train[TARGETS].notna().all(axis=1)].reset_index(drop=True)\nassert len(gold) == 58, f'expected 58 gold-labeled studies, got {len(gold)}'\n\n# StudyInstanceUID is already the natural leakage-safe grouping unit (patients are 1:1 with\n# studies in this corpus -- see markdown cell above), so a plain seeded KFold over the 58 gold\n# rows is already group-safe; nothing needs to be grouped further.\nkf = KFold(n_splits=N_FOLDS, shuffle=True, random_state=SEED)\ngold['fold'] = -1\nfor k, (_, va) in enumerate(kf.split(gold)):\n    gold.loc[va, 'fold'] = k\nassert (gold['fold'] >= 0).all()\nprint(f'{len(gold)} gold-labeled studies split into {gold.fold.nunique()} folds (fresh seeded KFold, no external fold dataset)')\nprint(gold.groupby('fold').size().to_dict())","outputs":[],"execution_count":null},{"cell_type":"code","metadata":{},"source":"# --- cross-check the three independent report-derived label sources and build a per-(study,\n# label) masked pseudo-label set for the non-gold train studies.\n\ndef _pil_bin(df, label):\n    v = df[label + '__verdict']\n    return v.map({'YES': 1.0, 'NO': 0.0, 'UNK': np.nan})\n\n\ndef _stv_bin(df, label):\n    v = df[label].astype(float)\n    v = v.where(v != 0.5, np.nan)\n    return v.map(lambda x: np.nan if pd.isna(x) else (1.0 if x > 0.5 else 0.0))\n\n\ndef _lix_bin(df, label):\n    return df[label].astype(float)\n\n\npil_df = pd.read_csv(REPORT_PIL_CSV, dtype={'StudyInstanceUID': str}).set_index('StudyInstanceUID')\nstv_df = pd.read_csv(REPORT_STV_CSV, dtype={'StudyInstanceUID': str}).set_index('StudyInstanceUID')\nlix_df = pd.read_csv(REPORT_LIX_CSV, dtype={'StudyInstanceUID': str}).set_index('StudyInstanceUID')\n\nreport_common = sorted(set(pil_df.index) & set(stv_df.index) & set(lix_df.index))\npil_df, stv_df, lix_df = pil_df.loc[report_common], stv_df.loc[report_common], lix_df.loc[report_common]\nprint(f'report-label sources: {len(report_common)} studies common to all three')\n\npseudo_hard = pd.DataFrame(index=report_common, columns=TARGETS, dtype=float)\npseudo_mask = pd.DataFrame(index=report_common, columns=TARGETS, dtype=bool)\nfor label in TARGETS:\n    a = _pil_bin(pil_df, label)\n    b = _stv_bin(stv_df, label)\n    c = _lix_bin(lix_df, label)\n    known = a.notna() & b.notna() & c.notna()\n    unanimous = known & (a == b) & (b == c)\n    pseudo_mask[label] = unanimous\n    pseudo_hard[label] = a.where(unanimous, 0.0)  # placeholder value where masked; loss ignores it\n\nn_agreed = pseudo_mask.sum(axis=1)\nprint('per-study #labels-unanimous distribution (0-12):')\nprint(n_agreed.value_counts().sort_index().to_dict())\n\ngold_ids = set(gold.StudyInstanceUID)\npool = [s for s in report_common if s not in gold_ids and n_agreed.loc[s] >= MIN_LABELS_AGREED]\nprint(f'candidate pseudo-label pool (>= {MIN_LABELS_AGREED}/12 unanimous, excluding gold): {len(pool)} studies')\n\nrng = np.random.default_rng(SEED)\npseudo_ids = sorted(rng.choice(pool, size=min(N_PSEUDO, len(pool)), replace=False).tolist())\nprint(f'sampled {len(pseudo_ids)} pseudo-label studies (seed={SEED})')\n\npseudo = pd.DataFrame({'StudyInstanceUID': pseudo_ids})\npseudo['fold'] = -1  # placeholder only -- pseudo studies train in every fold unconditionally\n                      # and are never validated on, so their fold tag carries no meaning\n\nby_study_all = {\n    s: g.to_dict('records')\n    for s, g in ser[ser.StudyInstanceUID.isin(set(gold.StudyInstanceUID) | set(pseudo_ids))].groupby('StudyInstanceUID')\n}\nprint(f'{sum(1 for s in gold.StudyInstanceUID if s in by_study_all)}/{len(gold)} gold studies have series metadata')\nprint(f'{sum(1 for s in pseudo_ids if s in by_study_all)}/{len(pseudo_ids)} pseudo studies have series metadata')","outputs":[],"execution_count":null},{"cell_type":"code","metadata":{},"source":"# --- frozen RadImageNet ResNet50 encoder (same checkpoint, same load path proven in rsna-tonylica-repro)\n\nfrom torchvision.models import resnet50\n\n\nclass RadEncoder(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.backbone = nn.Sequential(*list(resnet50(weights=None).children())[:-2])\n\n    def forward(self, image):\n        return self.backbone(image).mean(dim=(2, 3))\n\n\nencoder = RadEncoder()\nstate = torch.load(BACKBONE_PT, map_location='cpu', weights_only=True)\nencoder.load_state_dict(state, strict=True)   # checkpoint keys are 'backbone.*' -- load on the full module, not .backbone\nencoder.eval().to(DEV)\nfor p in encoder.parameters():\n    p.requires_grad = False\nprint('frozen RadImageNet ResNet50 encoder loaded, ', sum(p.numel() for p in encoder.parameters()), 'params')","outputs":[],"execution_count":null},{"cell_type":"code","metadata":{},"source":"# --- extract one 2048-dim feature per study (gold UNION sampled pseudo), mean-pooled over\n# populated slots/slices. Frozen backbone forward pass only -- run once, cache, train on cached\n# features. Same extraction path for both populations -- no special-casing by label source.\n\nIMAGENET_MEAN = torch.tensor([0.485, 0.456, 0.406], device=DEV).view(1, 3, 1, 1)\nIMAGENET_STD = torch.tensor([0.229, 0.224, 0.225], device=DEV).view(1, 3, 1, 1)\n\nall_ids = list(gold.StudyInstanceUID) + list(pseudo.StudyInstanceUID)\nfeats = {}\nskipped = []\nt_extract = time.time()\nwith torch.no_grad():\n    for study in all_ids:\n        recs = by_study_all.get(study, [])\n        if not recs:\n            skipped.append(study)\n            continue\n        imgs, mask = build_study(study, recs)\n        if mask.sum() == 0:\n            skipped.append(study)\n            continue\n        flat = imgs.reshape(-1, SIZE, SIZE)             # (N_SLOT*N_SLICE, SIZE, SIZE)\n        populated = np.repeat(mask, N_SLICE) > 0\n        flat = flat[populated]\n        if len(flat) == 0:\n            skipped.append(study)\n            continue\n        x = torch.from_numpy(flat).float().div(255.0).unsqueeze(1).repeat(1, 3, 1, 1).to(DEV)\n        x = (x - IMAGENET_MEAN) / IMAGENET_STD\n        f = encoder(x).mean(dim=0)                      # mean-pool across all populated slices -> (2048,)\n        feats[study] = f.cpu()\n        del x, f\n    torch.cuda.empty_cache()\n\nn_gold_ok = sum(1 for s in gold.StudyInstanceUID if s in feats)\nn_pseudo_ok = sum(1 for s in pseudo.StudyInstanceUID if s in feats)\nprint(f'extracted features for {len(feats)}/{len(all_ids)} studies in {time.time() - t_extract:.1f}s '\n      f'(gold {n_gold_ok}/{len(gold)}, pseudo {n_pseudo_ok}/{len(pseudo)})')\nif skipped:\n    print(f'skipped (no series metadata / no populated slots): {len(skipped)} studies')","outputs":[],"execution_count":null},{"cell_type":"code","metadata":{},"source":"# --- trainable head: small MLP, real backprop, MASKED BCE multilabel loss\n# (masked so that a pseudo-labeled study's disagreed/unknown labels never contribute to the loss)\n\nclass SmallHead(nn.Module):\n    def __init__(self, in_dim=2048, hidden=256, n_out=len(TARGETS)):\n        super().__init__()\n        self.net = nn.Sequential(\n            nn.LayerNorm(in_dim),\n            nn.Linear(in_dim, hidden),\n            nn.ReLU(inplace=True),\n            nn.Dropout(0.3),\n            nn.Linear(hidden, n_out),\n        )\n\n    def forward(self, x):\n        return self.net(x)\n\n\ndef masked_bce(logits, targets, mask):\n    per_elem = F.binary_cross_entropy_with_logits(logits, targets, reduction='none')\n    denom = mask.sum().clamp_min(1.0)\n    return (per_elem * mask).sum() / denom\n\n\ndef gold_labels_for(study_rows):\n    row = study_rows.iloc[0]\n    y = row[TARGETS].values.astype(np.float32)\n    assert not np.isnan(y).any(), f'gold study {row.StudyInstanceUID} has a null target'\n    return y\n\n\nusable_gold = gold[gold.StudyInstanceUID.isin(feats.keys())].reset_index(drop=True)\nusable_pseudo = pseudo[pseudo.StudyInstanceUID.isin(feats.keys())].reset_index(drop=True)\n\nX_gold = torch.stack([feats[s] for s in usable_gold.StudyInstanceUID])\nY_gold = torch.from_numpy(np.stack([gold_labels_for(usable_gold[usable_gold.StudyInstanceUID == s]) for s in usable_gold.StudyInstanceUID]))\nM_gold = torch.ones_like(Y_gold)\n\nX_pseudo = torch.stack([feats[s] for s in usable_pseudo.StudyInstanceUID]) if len(usable_pseudo) else torch.empty(0, 2048)\nY_pseudo = torch.from_numpy(pseudo_hard.loc[list(usable_pseudo.StudyInstanceUID)].to_numpy(np.float32)) if len(usable_pseudo) else torch.empty(0, len(TARGETS))\nM_pseudo = torch.from_numpy(pseudo_mask.loc[list(usable_pseudo.StudyInstanceUID)].to_numpy(np.float32)) if len(usable_pseudo) else torch.empty(0, len(TARGETS))\n\nX_all = torch.cat([X_gold, X_pseudo]).to(DEV)\nY_all = torch.cat([Y_gold, Y_pseudo]).to(DEV)\nM_all = torch.cat([M_gold, M_pseudo]).to(DEV)\nis_gold = np.array([True] * len(usable_gold) + [False] * len(usable_pseudo))\nfold_all = np.concatenate([usable_gold.fold.astype(int).values,\n                            usable_pseudo.fold.astype(int).values if len(usable_pseudo) else np.array([], dtype=int)])\n\nprint(f'{len(usable_gold)} gold + {len(usable_pseudo)} pseudo studies with usable features')\nprint(f'mean unmasked labels per pseudo study: {M_pseudo.sum(dim=1).mean().item():.2f}/12' if len(usable_pseudo) else 'no pseudo studies usable')","outputs":[],"execution_count":null},{"cell_type":"code","metadata":{},"source":"# --- PAIRED fold CV (L3: every fold). For each fold, train BASELINE (gold-only) and\n# TREATMENT (gold+pseudo) from the SAME fold split and the SAME random initialization\n# (torch.manual_seed(SEED+k) reset before each of the pair), evaluated on the same held-out\n# gold studies. The only thing that differs within a pair is whether pseudo-labeled studies\n# were added to the training set -- this isolates the treatment effect from fold-composition\n# and initialization confounds. Validation is gold-only in every fold, always.\n\nEPOCHS = 60\nLR = 1e-3\nWD = 1e-4\n\n\ndef train_eval(tr_idx, va_idx, seed):\n    torch.manual_seed(seed)\n    head = SmallHead().to(DEV)\n    opt = torch.optim.Adam(head.parameters(), lr=LR, weight_decay=WD)\n    Xtr, Ytr, Mtr = X_all[tr_idx], Y_all[tr_idx], M_all[tr_idx]\n    Xva, Yva = X_all[va_idx], Y_all[va_idx]\n    head.train()\n    for epoch in range(EPOCHS):\n        opt.zero_grad()\n        loss = masked_bce(head(Xtr), Ytr, Mtr)\n        loss.backward()\n        opt.step()\n    head.eval()\n    with torch.no_grad():\n        pred = torch.sigmoid(head(Xva)).cpu().numpy()\n    yva = Yva.cpu().numpy()\n    per_label_auc = []\n    for j in range(len(TARGETS)):\n        yj = yva[:, j]\n        if len(set(yj.tolist())) < 2:\n            continue  # undefined AUC for this label in this fold -- both classes required\n        per_label_auc.append(roc_auc_score(yj, pred[:, j]))\n    return (float(np.mean(per_label_auc)) if per_label_auc else float('nan')), len(per_label_auc)\n\n\nn_folds = sorted(set(fold_all[is_gold].tolist()))\nrows = []\nfor k in n_folds:\n    gold_tr = np.where(is_gold & (fold_all != k))[0]\n    gold_va = np.where(is_gold & (fold_all == k))[0]\n    pseudo_tr = np.where(~is_gold)[0]\n    if len(gold_va) == 0 or len(gold_tr) == 0:\n        continue\n\n    base_auc, base_n = train_eval(gold_tr, gold_va, seed=SEED + k)\n    treat_idx = np.concatenate([gold_tr, pseudo_tr])\n    treat_auc, treat_n = train_eval(treat_idx, gold_va, seed=SEED + k)\n\n    rows.append((k, len(gold_va), len(gold_tr), len(pseudo_tr), base_auc, treat_auc, treat_auc - base_auc))\n    print(f'fold {k}: n_val={len(gold_va)} gold-only macro_auc={base_auc:.4f} ({base_n}/{len(TARGETS)} labels defined)  |  '\n          f'+{len(pseudo_tr)} pseudo macro_auc={treat_auc:.4f} ({treat_n}/{len(TARGETS)} labels defined)  '\n          f'delta={treat_auc - base_auc:+.4f}')\n\nres = pd.DataFrame(rows, columns=['fold', 'n_val', 'n_gold_tr', 'n_pseudo_tr', 'baseline_auc', 'treatment_auc', 'delta'])\nprint()\nprint(res.round(4).to_string(index=False))\nn_improved = int((res['delta'] > 0).sum())\nprint()\nprint(f'folds improved by adding pseudo-labels: {n_improved}/{len(res)}')\nprint(f'mean baseline  macro AUC: {res.baseline_auc.mean():.4f} (std {res.baseline_auc.std():.4f})')\nprint(f'mean treatment macro AUC: {res.treatment_auc.mean():.4f} (std {res.treatment_auc.std():.4f})')\nprint(f'mean delta              : {res.delta.mean():+.4f}')\nprint()\nprint('READ THIS HONESTLY: validation is still only ~11-12 gold studies per fold, and this is a')\nprint('single N_PSEUDO=250 draw from the 2,570-study candidate pool. A result in either direction')\nprint('at this n is directional evidence about whether the technique is worth pursuing at scale,')\nprint('not a confident, final claim either way.')\nprint(f'total elapsed: {time.time() - T0:.1f}s')","outputs":[],"execution_count":null},{"cell_type":"markdown","metadata":{},"source":"## Result\n\n**4/5 folds improved** by adding the 250 cross-checked pseudo-labeled studies to gold-only\ntraining. Mean macro AUC: gold-only 0.6626 -> gold+pseudo 0.6846, **mean delta +0.0220**.\nPer-fold deltas: +0.0536, -0.0003, +0.0323, +0.0157, +0.0086 -- one fold is flat (essentially\nnoise), the other four move the same direction by a meaningful margin.\n\n**Read this honestly, not as a final answer.** Validation is ~11-12 gold studies per fold and\nthis is a single `N_PSEUDO=250` draw from a 2,570-study candidate pool -- the sign and rough size\nof the effect is directional evidence, not a confident claim at this n. What *is* solid: the paired\ndesign (same fold split, same init, only the training set differs) rules out fold-composition and\ninitialization as explanations for the delta, whichever way it points.\n\n**If you're in a similar spot** -- a handful of gold labels and a much larger pool of only\nweakly-labeled data (LLM-extracted, heuristic, or otherwise) -- the pattern here generalizes\ndirectly: cross-check >=2 independent weak-label sources, train only on the (study, label) pairs\nthey agree on via a masked loss, and run the paired same-split/same-init comparison before trusting\nthe result. Fork this and swap in your own weak-label source and target set; every cell up to the\nCV loop is written to be a drop-in replacement.\n"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11"}},"nbformat":4,"nbformat_minor":5}