{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Competition Overview\nThe knee is the most commonly injured and imaged joint in the body, however, the ways in which radiologists interpret MRI imaging scans differ. The goal of the competition is to develop ML models that detect clinically important knee abnormalities, which could provide useful decision support tools to radiologists in practice.\n\n## Competition Training Data & Prior Findings\n- Basic data inventory, EDA, and image processing exploration in our notebooks [RSNA_Data_Inventory](https://www.kaggle.com/code/joshuaziel/rsna-data-inventory),   \n  [RSNA_EDA_Preliminary](https://www.kaggle.com/code/joshuaziel/rsna-eda-preliminary), [RSNA_Image_Processing_Exploration](https://www.kaggle.com/code/joshuaziel/rsna-image-processing-exploration), respectively.  Many other great notebooks are out there on the same topics.  We labeled the pre-labeled languages for each report in [RSNA_label_report_languages](https://www.kaggle.com/code/joshuaziel/rsna-label-report-languages)\n- The training data itself consists of over 4000 studies, each containing multiple MRI series. Only 58 training studies are labeled, however, radiologic reports are available \n  for all of the training data and may be used to psuedolabel all or some of the remaining training studies\n- Clear differences between the gold labels and the reports are expected: a [conversation (733826)](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection/discussion/733826) \n  on this precise topic, where the host confirmed independent reads.  The conversations examples align with this experience and helpfully, there's a [conversation (733343)](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection/discussion/733343) which includes the definitions used for labeling.\n- LLMs have been used increasingly for generation, summarization, and classification tasks related to radiologic reports ([Lee RC et al, 2026](https://www.jacr.org/article/S1546-1440(25)00584-8/fulltext), [Abdullah A et al, 2025](https://pmc.ncbi.nlm.nih.gov/articles/PMC11970564/)).  Our plan is to take a similar approach and use an LLM with an optimized prompt to generate pseudolabels for the unlabeled training studies\n- We've previously explored in [RSNA_Slightly_Overexhaustive_Labeling_Deathmatch](https://www.kaggle.com/code/joshuaziel/rsna-slightly-overexhaustive-labeling-deathmatch). \n\n## Goal for Notebook: \nLeverage the dspy 'Consensus' program from the prior notebook and generate a comprehensive set of predicted labels for the report using OpenAI's gpt5.6 'luna' model.\n","metadata":{}},{"cell_type":"code","source":"!pip install dspy\nimport pandas as pd\nimport numpy as np\nimport asyncio\nimport shutil\nimport aiofiles\nimport os\nimport json\nimport dspy\nfrom sklearn.metrics import balanced_accuracy_score\nfrom typing import TypedDict, Literal, Optional, Annotated\nfrom pydantic import BaseModel, Field\nfrom pathlib import Path\nfrom kaggle_secrets import UserSecretsClient\nfrom dotenv import load_dotenv, find_dotenv\nfrom enum import Enum\nfrom tqdm import tqdm\n\nuser_secrets = UserSecretsClient()\nOPENAI_API_KEY = user_secrets.get_secret(\"OPENAI_API_KEY\")\nprint(\"Successfully loaded OPENAI_API_KEY.\")\n\nDATA_LOC = Path('/kaggle/input/competitions/rsna-knee-abnormality-detection')\nDATASET_LOC = Path('/kaggle/input/datasets/joshuaziel/rsna-knee-mri-report-labels-gpt-5-6-luna/all_raw_predictions.json')\nOUTPUTS_LOC = Path('/kaggle/working')\n\nTRAIN_COLUMNS = ['StudyInstanceUID', 'Report', 'ACL', 'MCL', 'Medial_Meniscus', 'Lateral_Meniscus',\n                 'Medial_OA', 'Lateral_OA', 'PF_OA', 'Effusion', 'Synovitis', 'Baker', \n                 'Contusion', 'Fracture']\n\nFINDING_COLUMNS = ['ACL', 'MCL', 'Medial_Meniscus', 'Lateral_Meniscus',\n                 'Medial_OA', 'Lateral_OA', 'PF_OA', 'Effusion', 'Synovitis', 'Baker', \n                 'Contusion', 'Fracture']\n\ntrain_df = pd.read_csv(DATA_LOC / 'train.csv', names = TRAIN_COLUMNS, header = 0)\nprint(f\"Successfully loaded {train_df.shape[0]} rows from train.csv.\")\n\ndspy_lm = dspy.LM(model = 'openai/gpt-5.6-luna', temperature = 1.0, model_type = 'chat', api_key = OPENAI_API_KEY, cache = False)        \ndspy.configure(lm=dspy_lm)\ndspy.settings.configure(track_usage=True) \nprint(f\"DSPy language model set to {dspy.settings.lm.model}.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:33:11.272068Z","iopub.execute_input":"2026-09-18T20:33:11.27242Z","iopub.status.idle":"2026-09-18T20:33:15.683752Z","shell.execute_reply.started":"2026-09-18T20:33:11.27239Z","shell.execute_reply":"2026-09-18T20:33:15.682792Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class ClassificationLabelConsensus(Enum):\n    PRESENT = 1\n    ABSENT = 0\n\n\nclass ConfidenceRatingConsensus(Enum):\n    VERY_HIGH = 5\n    HIGH = 4\n    MODERATE = 3\n    LOW = 2\n    VERY_LOW = 1\n    NO_CONFIDENCE = 0\n\n\nclass EvidenceBasisConsensus(Enum):\n    \"\"\"The kind of report evidence the annotation rests on. Independent of the\n    annotation: STATED and DESCRIBED can each support PRESENT or ABSENT.\"\"\"\n\n    STATED = \"the report names the finding and either asserts or denies it\"\n    DESCRIBED = \"the finding is not named; the annotation follows from described morphology, or from the relevant structure being described as normal\"\n    NOT_MENTIONED = \"the report does not mention this finding\"\n\n\nclass FindingAnnotationConsensus(TypedDict):\n    annotation: Annotated[\n        ClassificationLabelConsensus,\n        Field(\n            description=\n\"\"\"PRESENT or ABSENT for this finding, per the finding definition. Every class takes one of the\ntwo; there is no abstention.\"\"\"\n        ),\n    ]\n    confidence: Annotated[\n        ConfidenceRatingConsensus,\n        Field(\n            description=\n\"\"\"Confidence in the specific annotation made based on the annotation policy\"\"\"\n        ),\n    ]\n    evidence_basis: Annotated[\n        EvidenceBasisConsensus,\n        Field(\n            description=\n\"\"\"The kind of report evidence this annotation rests on. Select the most specific applicable\ncategory: prefer STATED or DESCRIBED, and use NOT_MENTIONED only when neither applies.\"\"\"\n        ),\n    ]\n    rationale: Annotated[\n        str,\n        Field(\n            description=\n\"\"\"Short justification (in English, regardless of the report language) for the assigned\nannotation and confidence. Quote or closely paraphrase the report language relied on, or state\nthat the report is silent.\"\"\"\n        ),\n    ]\n\n\nclass ReportFindingsConsensus(TypedDict):\n    ACL: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A high-grade partial or full-thickness tear of the anterior cruciate ligament, meaning complete discontinuity\nof the ligament, or more than 50 percent of fibers disrupted, with or without secondary signs such as characteristic\npivot-shift bone contusions. Mild signal change, degeneration, or thickening without discontinuity is graded negative.\"\"\"\n        ),\n    ]\n    MCL: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A high-grade partial or complete acute tear of the medial collateral ligament, with disrupted fibers\nand edema within and adjacent to the ligament. Low-grade sprains and chronic or remote stress changes are graded negative.\"\"\"\n        ),\n    ]\n    Medial_Meniscus: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the medial meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Lateral_Meniscus: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the lateral meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Medial_OA: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the medial compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Lateral_OA: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the lateral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    PF_OA: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the patellofemoral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Effusion: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large amount of fluid distending the joint.\"\"\"\n        ),\n    ]\n    Synovitis: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"Inflammation and thickening of the synovial lining of the joint.\"\"\"\n        ),\n    ]\n    Baker: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large fluid collection in the characteristic Baker (popliteal) cyst location behind the knee.\"\"\"\n        ),\n    ]\n    Contusion: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A bone contusion, seen as bone marrow edema-like signal from impact, without a discrete fracture line.\"\"\"\n        ),\n    ]\n    Fracture: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"An acute cortical break or fracture line.\"\"\"\n        ),\n    ]\n\n\n# Same 12 names as FINDING_COLUMNS, derived from the schema so the two cannot drift apart\nFINDING_NAMES_CONSENSUS = tuple(ReportFindingsConsensus.__annotations__)\nENUM_KEYS_CONSENSUS = {'annotation': ClassificationLabelConsensus, 'confidence': ConfidenceRatingConsensus}\n\n\ndef _as_int_consensus(value, enum_cls: type[Enum]) -> int | None:\n    if value is None:\n        return None\n    if isinstance(value, enum_cls):\n        return int(value.value)\n    if isinstance(value, str):\n        name = value.strip().upper()\n        if name in enum_cls.__members__:\n            return int(enum_cls[name].value)\n        value = int(value)\n    return int(enum_cls(int(value)).value)\n\n\ndef _as_evidence_basis_name_consensus(value) -> str | None:\n    if value is None:\n        return None\n    if isinstance(value, EvidenceBasisConsensus):\n        return value.name\n    name = str(value).strip().upper()\n    if name in EvidenceBasisConsensus.__members__:\n        return name\n    return EvidenceBasisConsensus(str(value)).name\n\n\ndef as_long_records_consensus(findings: dict, study_instance_uid: str) -> list[dict]:\n    records = []\n    for name in FINDING_NAMES_CONSENSUS:\n        annotation = findings.get(name) or {}\n        rationale = annotation.get('rationale')\n        records.append({\n            'StudyInstanceUID': str(study_instance_uid),\n            'Finding': name,\n            **{key: _as_int_consensus(annotation.get(key), cls) for key, cls in ENUM_KEYS_CONSENSUS.items()},\n            'evidence_basis': _as_evidence_basis_name_consensus(annotation.get('evidence_basis')),\n            'rationale': None if rationale is None else str(rationale),\n        })\n    return records\n\n\nclass KneeFindingsConsensus(BaseModel):\n    findings: ReportFindingsConsensus = Field(\n        description=\n\"\"\"Annotations per class for the knee MRI radiology report\"\"\"\n    )\n\n\nclass KneeMRIOutputConsensus(dspy.Signature):\n    \"\"\"\n    Act as an expert musculoskeletal radiologist annotating a knee MRI radiology report for\n    each class of finding. Use the definition supplied with each class when making annotations.\n\n    DECISION PROCEDURE\n    For each class determine in order whether the report text provides:\n      1. A clear and direct statement confirming the presence of the finding (or its absence)\n         according to the finding criteria, including any severity thresholds?\n      2. Equivalent language describing the finding's presence (or absence) (eg, using clinically\n         related language or as part of a composite diagnosis)?\n      3. A normal or negative statement about the relevant structure or compartment that would\n         indicate the finding is absent (eg, the relevant structure is described as normal or\n         without defects)?\n      4. Another reported finding that makes this one likely or unlikely, given normal co-occurrence\n         of knee pathology?\n      5. Silence on the finding? If none of the above applies, annotate ABSENT.\n    Give the impression/conclusion precedence over the body when they conflict, and say so in\n    the rationale.\n\n    ANNOTATION POLICY\n    Every class receives PRESENT or ABSENT; there is no abstention. A report that is silent,\n    equivocal, or non-diagnostic for a class still receives the more likely of the two labels, with\n    the residual uncertainty carried by the confidence rating rather than by withholding a label.\n    - PRESENT: the report affirms the finding, describes its substance, or supports it strongly\n      enough that an independent image reader would more likely than not have recorded it.\n    - ABSENT: the report denies the finding, describes the relevant structures as normal, or is\n      silent about it.\n\n    HOW TO HANDLE SILENCE\n    - Silence resolves to ABSENT.\n    - Silence is weaker evidence of absence for findings that are frequently observed but omitted as\n      incidental (synovial change, compartment-specific cartilage grading in a trauma report).\n      Annotate ABSENT in those cases as well, and carry the weakness in a lower confidence rating.\n\n    HOW TO HANDLE SEVERITY QUALIFIERS\n      - Where a class definition has a size or severity threshold, do not resolve to ABSENT on the\n        strength of a diminutive adjective alone when the finding itself is affirmed. Weigh the\n        objective content of the description (measurements, extent, number of regions, secondary\n        signs) more importantly.\n      - Distinguish a diminutive qualifier from a negation. \"Mild effusion\" affirms the finding;\n        \"no significant effusion\" and \"no effusion\" indicate the finding is ABSENT.\n      - A finding that is named but not quantified should generally be graded PRESENT, as\n        radiologists tend to name findings that are conspicuous.\n      - The exception is the cartilage/osteoarthritis classes, where the grading vocabulary is\n        standardized (Outerbridge/ICRS grade, \"full-thickness\", \"advanced\", \"complete cartilage\n        loss\") and does track the reference threshold. There, respect the reported grade, and\n        treat low-grade chondral change, fissuring, thinning, an unqualified diagnosis of\n        \"osteoarthritis\", or isolated osteophytes as not meeting the definition.\n\n    FINDINGS ARE NOT MUTUALLY EXCLUSIVE\n    Annotate each class independently against its own definition. In particular, a report that\n    attributes an abnormality to one entity does not thereby exclude an overlapping class: bone\n    marrow edema accompanying a fracture still satisfies a marrow-edema definition, and\n    traumatic and degenerative explanations for the same signal can both be recorded.\n\n    CONFIDENCE\n    Confidence is the confidence in the specific annotation according to the policy above; it is\n    NOT confidence that a finding is present. Annotations of PRESENT and of ABSENT will be made\n    with varying levels of confidence, which is what the rating should reflect. An ABSENT\n    annotation resting on silence alone, or a call that turns on an unquantified severity\n    qualifier, belongs at the low end of the scale.\n\n    EVIDENCE BASIS\n    Commit to the kind of report evidence the annotation rests on before writing the rationale.\n    STATED and DESCRIBED each support PRESENT or ABSENT; NOT_MENTIONED records that the report is\n    silent, and pairs with an ABSENT annotation under the silence policy.\n\n    RATIONALE\n    The rationale must quote or closely paraphrase the specific report language relied on, or state\n    that the report is silent. It must be consistent with the evidence basis and with the annotation.\n    \"\"\"\n    report: str = dspy.InputField()\n    classification: KneeFindingsConsensus = dspy.OutputField()\n\n\nclass ConsensusClassifier(dspy.Module):\n    def __init__(self):\n        super().__init__()\n        self.classification_module = dspy.ChainOfThought(KneeMRIOutputConsensus)\n\n    def forward(self, report: str, study_instance_uid: str) -> dspy.Prediction:\n        result = self.classification_module(report=report)\n\n        return dspy.Prediction(\n            study_instance_uid=study_instance_uid,\n            findings=as_long_records_consensus(result.classification.findings, study_instance_uid),\n        )\n\nprint(\"Defined DSPy program elements.\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:33:18.16282Z","iopub.execute_input":"2026-09-18T20:33:18.1638Z","iopub.status.idle":"2026-09-18T20:33:18.194555Z","shell.execute_reply.started":"2026-09-18T20:33:18.16376Z","shell.execute_reply":"2026-09-18T20:33:18.193589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"async def asave_records(filepath: str, records: list[dict]):\n    json_records = json.dumps(records, indent = 4)\n    async with aiofiles.open(filepath, mode='w') as f:\n        await f.write(json_records)\n\nasync def apredict_record(record:dict, semaphore: asyncio.BoundedSemaphore, program:dspy.Module) -> list[dict]:\n    async with semaphore:\n        study_instance_uid = record.get(\"StudyInstanceUID\")\n        report = record.get(\"Report\")\n        print(f'Getting predictions for {study_instance_uid}')\n        if (study_instance_uid is not None) and (report is not None):\n            success, tries = False, 0\n            while not success and tries < 3:\n                tries += 1\n                try: \n                    prediction = await program(study_instance_uid = study_instance_uid, report = report)\n                    success = True\n                    return prediction.findings\n                except Exception as e:\n                    print(f'Error prediction labels for {study_instance_uid} on try {tries} of 3:{e}')\n            if not success:\n                return as_long_records_consensus({}, study_instance_uid)\n                print(f'Unable to predict labels for {study_instance_uid} due to repeated errors.')\n        else:\n            print('Unable to predict labels: could not recover StudyInstanceUID, Report, or both')\n            return as_long_records_consensus({}, study_instance_uid)\n\nasync def apredict_batch(batch:list[dict], semaphore: asyncio.BoundedSemaphore, program:dspy.Module) -> list[dict]:\n    batch_items = [apredict_record(x, semaphore, program) for x in batch]\n    results = await asyncio.gather(*batch_items)\n    return [finding for findings in results for finding in findings]\n\nasync def apredict_all_records(records:pd.DataFrame, sem:int, program:dspy.Module, batch_size:int = 100, outputs_path:Path = OUTPUTS_LOC, name_base:str|None = None) -> list[dict]:\n    def make_batches(records: pd.DataFrame, batch_size:int) -> list[list[dict]]:\n        if records.shape[0] < batch_size:\n            batch_count = 1\n        else:\n            batch_count = (int(records.shape[0]/batch_size)+ 1) if int(records.shape[0]/batch_size) < (records.shape[0]/batch_size) else int(records.shape[0] / batch_size)\n        batches =[records[['StudyInstanceUID', 'Report']].iloc[(i * batch_size):(i*batch_size + batch_size), :].to_dict(orient = 'records') for i in range(batch_count)]\n        print(f'Created {len(batches)} batches.')\n        return batches\n    \n    semaphore = asyncio.BoundedSemaphore(sem)\n    batches = make_batches(records, batch_size)\n    predictions = []\n    program = dspy.asyncify(program())\n    for i, batch in enumerate(batches):\n        batch_name = f\"batch_{i+1}\" if name_base is None else f\"{name_base}_batch_{i}\"\n        print(f'Running predictions for batch {i + 1} of {len(batches)}.')\n        result = await apredict_batch(batch, semaphore, program)\n        cache_path = outputs_path / f\"cache_{batch_name}.json\"\n        await asave_records(cache_path, result)\n        print(f\"Batch #{i+1} cached at {str(cache_path)}\")\n        predictions.extend(result)\n    results_name = \"all_raw_predictions.json\" if name_base is None else f\"{name_base}_all_raw_predictions.json\"\n    results_path = outputs_path / results_name\n    await asave_records(results_path, predictions)\n    print(f\"All records saved at {str(results_path)}\")\n    return predictions\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:33:27.298895Z","iopub.execute_input":"2026-09-18T20:33:27.299247Z","iopub.status.idle":"2026-09-18T20:33:27.315245Z","shell.execute_reply.started":"2026-09-18T20:33:27.299196Z","shell.execute_reply":"2026-09-18T20:33:27.314465Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"if DATASET_LOC.is_file():\n    with open(DATASET_LOC, 'r') as f:\n        all_raw_predictions = json.load(f)\n    print(f'Loaded LLM label predictions from file: {len(all_raw_predictions)} records (from {int((len(all_raw_predictions))/12)} studies in \"train.csv\").')\n\nelse:\n    all_raw_predictions = await apredict_all_records(records = train_df, program = ConsensusClassifier, batch_size = 500, sem = 50)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:33:31.409175Z","iopub.execute_input":"2026-09-18T20:33:31.409542Z","iopub.status.idle":"2026-09-18T20:33:31.584496Z","shell.execute_reply.started":"2026-09-18T20:33:31.409511Z","shell.execute_reply":"2026-09-18T20:33:31.583547Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"all_raw_predictions_df = pd.DataFrame(all_raw_predictions)\nall_raw_predictions_df.head(12)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:33:34.453261Z","iopub.execute_input":"2026-09-18T20:33:34.453615Z","iopub.status.idle":"2026-09-18T20:33:34.548036Z","shell.execute_reply.started":"2026-09-18T20:33:34.453584Z","shell.execute_reply":"2026-09-18T20:33:34.547252Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"assert sum(pd.isna(all_raw_predictions)) == 0, 'Unexpected missing values detected'\n\ngold_labeled_ids = train_df.dropna()['StudyInstanceUID']\nassert (len(all_raw_predictions_df[all_raw_predictions_df['StudyInstanceUID'].isin(gold_labeled_ids)]['StudyInstanceUID'].unique()) == len(gold_labeled_ids)), 'Gold-labeled sudies from predictions are missing'\n\nprint(f\"Original dataset contains {len(train_df['StudyInstanceUID'])} studies: long data contains:\")\nall_raw_predictions_df.groupby(['StudyInstanceUID']).count().count()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:33:37.315962Z","iopub.execute_input":"2026-09-18T20:33:37.316316Z","iopub.status.idle":"2026-09-18T20:33:37.378416Z","shell.execute_reply.started":"2026-09-18T20:33:37.316283Z","shell.execute_reply":"2026-09-18T20:33:37.377725Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"long_gold_df = (train_df.dropna()\n        .melt(id_vars = ['StudyInstanceUID', 'Report'], var_name = 'Finding', value_name = 'gold_label')\n        .drop(columns = 'Report'))\nlong_gold_df['gold_label'] =  long_gold_df['gold_label'].astype('Int64') \n\nlong_gold_predictions_df = pd.merge(\n    all_raw_predictions_df,\n    long_gold_df,\n    on = ['StudyInstanceUID', 'Finding'],\n    how = 'inner'\n)\n\nlong_gold_predictions_df[long_gold_predictions_df['StudyInstanceUID'].isin(gold_labeled_ids)].head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:35:07.811654Z","iopub.execute_input":"2026-09-18T20:35:07.812005Z","iopub.status.idle":"2026-09-18T20:35:07.856356Z","shell.execute_reply.started":"2026-09-18T20:35:07.811974Z","shell.execute_reply":"2026-09-18T20:35:07.855544Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"unique_studies = long_gold_predictions_df['StudyInstanceUID'].unique()\nba_score = balanced_accuracy_score(long_gold_predictions_df['gold_label'], long_gold_predictions_df['annotation'])\nprint(f\"Balanced accuracy across all labels for {len(unique_studies)} studies: {round(ba_score, ndigits = 3)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:37:50.123866Z","iopub.execute_input":"2026-09-18T20:37:50.124305Z","iopub.status.idle":"2026-09-18T20:37:50.135844Z","shell.execute_reply.started":"2026-09-18T20:37:50.124265Z","shell.execute_reply":"2026-09-18T20:37:50.134954Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"all_raw_predictions_wide_df = (\n    all_raw_predictions_df\n        .drop(columns = ['confidence', 'evidence_basis', 'rationale'])\n        .pivot(index = ['StudyInstanceUID'], columns = 'Finding', values = 'annotation')\n)\nall_raw_predictions_wide_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-18T20:02:25.580906Z","iopub.execute_input":"2026-09-18T20:02:25.581295Z","iopub.status.idle":"2026-09-18T20:02:25.627253Z","shell.execute_reply.started":"2026-09-18T20:02:25.581258Z","shell.execute_reply":"2026-09-18T20:02:25.626448Z"}},"outputs":[],"execution_count":null}]}