{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.12.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Competition Overview\nThe knee is the most commonly injured and imaged joint in the body, however, the ways in which radiologists interpret MRI imaging scans differ. The goal of the competition is to develop ML models that detect clinically important knee abnormalities, which could provide useful decision support tools to radiologists in practice.\n\n## Competition Training Data & Prior Findings\n- Basic data inventory, EDA, and image processing exploration in our notebooks [RSNA_Data_Inventory](https://www.kaggle.com/code/joshuaziel/rsna-data-inventory),   \n  [RSNA_EDA_Preliminary](https://www.kaggle.com/code/joshuaziel/rsna-eda-preliminary), [RSNA_Image_Processing_Exploration](https://www.kaggle.com/code/joshuaziel/rsna-image-processing-exploration), respectively.  Many other great notebooks are out there on the same topics.  We labeled the pre-labeled languages for each report in [RSNA_label_report_languages](https://www.kaggle.com/code/joshuaziel/rsna-label-report-languages)\n- The training data itself consists of over 4000 studies, each containing multiple MRI series. Only 58 training studies are labeled, however, radiologic reports are available \n  for all of the training data and may be used to psuedolabel all or some of the remaining training studies\n- Clear differences between the gold labels and the reports are expected: a [conversation (733826)](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection/discussion/733826) \n  on this precise topic, where the host confirmed independent reads.  The conversations examples align with this experience and helpfully, there's a [conversation (733343)](https://www.kaggle.com/competitions/rsna-knee-abnormality-detection/discussion/733343) which includes the definitions used for labeling.\n- LLMs have been used increasingly for generation, summarization, and classification tasks related to radiologic reports ([Lee RC et al, 2026](https://www.jacr.org/article/S1546-1440(25)00584-8/fulltext), [Abdullah A et al, 2025](https://pmc.ncbi.nlm.nih.gov/articles/PMC11970564/)).  Our plan is to take a similar approach and use an LLM with an optimized prompt to generate pseudolabels for the unlabeled training studies\n\n## Goals for Notebook: \n1) Evaluate performance of LLMs from Anthropic, OpenAI on classification according to the 12 labels\n2) Select and to the extent possible optimize instructions/model combination or combinations for pseudolabeling\n","metadata":{}},{"cell_type":"markdown","source":"## 1. Load Imports and Data\n\n### Key Packages\n- Dspy for most LLM interactions and prompt optimizations; openai SDK for the rest\n- Pydantic for structured outputs\n- Plotnine for visualization \n\n### Key Data \n- Competition train.csv\n- `train_lang_lbl.csv` (Report langauges labeled from [RSNA_label_report_languages](https://www.kaggle.com/code/joshuaziel/rsna-label-report-languages))\n- Note that throughout the notebook, certain steps are serialized to json or txt to prevent re-running long loops; if you fork and use in kaggle,\n  either take the outputs too, or save such outputs to the inputs\n","metadata":{}},{"cell_type":"code","source":"!pip install dspy\nimport pandas as pd\nimport numpy as np\nimport os\nimport json\nimport dspy\nimport phik\nfrom typing import TypedDict, Literal, Optional, Annotated\nfrom openai import OpenAI\nfrom pydantic import BaseModel, Field\nfrom pathlib import Path\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import balanced_accuracy_score, cohen_kappa_score, f1_score, precision_score, recall_score, confusion_matrix, roc_auc_score, roc_curve\nfrom sklearn.linear_model import LogisticRegression\nfrom collections.abc import Callable\nfrom collections import namedtuple\nfrom dspy.evaluate.evaluate import EvaluationResult\nfrom kaggle_secrets import UserSecretsClient\nfrom dotenv import load_dotenv, find_dotenv\nfrom enum import Enum\nfrom tqdm import tqdm\nfrom plotnine import *\n\nuser_secrets = UserSecretsClient()\n#load_dotenv(find_dotenv())\n\n#OPENAI_API_KEY = os.getenv('OPENAI_API_KEY')\nOPENAI_API_KEY = user_secrets.get_secret(\"OPENAI_API_KEY\")\n#ANTHROPIC_API_KEY = os.getenv('ANTHROPIC_API_KEY')\nANTHROPIC_API_KEY = user_secrets.get_secret(\"ANTRHOPIC_API_KEY\")\n\nModelSpec = namedtuple('ModelSpec', ['name', 'key', 'type'])\n\nMODELS = {\n    'nano': ModelSpec('openai/gpt-5.4-nano-2026-03-17', OPENAI_API_KEY, 'chat'),\n    'luna': ModelSpec('openai/gpt-5.6-luna',  OPENAI_API_KEY, 'chat'),\n    'terra': ModelSpec('openai/gpt-5.6-terra',  OPENAI_API_KEY, 'chat'),\n    'sol': ModelSpec('openai/gpt-5.6-sol',  OPENAI_API_KEY, 'chat'),\n    'haiku': ModelSpec('anthropic/claude-haiku-4-5-20251001', ANTHROPIC_API_KEY, 'chat'),\n    'sonnet': ModelSpec('anthropic/claude-sonnet-5', ANTHROPIC_API_KEY, 'chat'),\n    'opus': ModelSpec('anthropic/claude-opus-5', ANTHROPIC_API_KEY, 'chat'),\n    'fable': ModelSpec('anthropic/claude-fable-5-1', ANTHROPIC_API_KEY, 'chat')\n}\n\nMODEL_KEYS = list(MODELS)\n\nLANG_GROUPS = ['English', 'Non-English/Unknown']\n\nDATA_LOC = Path('/kaggle/input')\nOUTPUTS_LOC = Path('/kaggle/working')\n\nTRAIN_COLUMNS = ['StudyInstanceUID', 'Report', 'ACL', 'MCL', 'Medial_Meniscus', 'Lateral_Meniscus',\n                 'Medial_OA', 'Lateral_OA', 'PF_OA', 'Effusion', 'Synovitis', 'Baker', \n                 'Contusion', 'Fracture']\n\nFINDING_COLUMNS = ['ACL', 'MCL', 'Medial_Meniscus', 'Lateral_Meniscus',\n                 'Medial_OA', 'Lateral_OA', 'PF_OA', 'Effusion', 'Synovitis', 'Baker', \n                 'Contusion', 'Fracture']\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:13:59.766503Z","iopub.execute_input":"2026-09-17T20:13:59.766932Z","iopub.status.idle":"2026-09-17T20:14:05.212682Z","shell.execute_reply.started":"2026-09-17T20:13:59.766887Z","shell.execute_reply":"2026-09-17T20:14:05.211511Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df = pd.read_csv(DATA_LOC / 'competitions' / 'rsna-knee-abnormality-detection' / 'train.csv', names = TRAIN_COLUMNS, header = 0)\nreport_langs_df = pd.read_csv(DATA_LOC / 'notebooks' / 'joshuaziel' / 'rsna-label-report-languages' / 'train_lang_lbl.csv', index_col = 0)\ngold_df = train_df.dropna().merge(report_langs_df, on = 'StudyInstanceUID', how = 'left')\ngold_df['Is_English'] = gold_df['Language'].apply(lambda x: 'English' if x == 'en' else 'Non-English/Unknown')\ngold_df['Num_Findings'] = gold_df.iloc[:,2:14].sum(axis = 1)\ngold_df['Above_Median_Findings'] = gold_df['Num_Findings'].apply(lambda x: x > 4)\ngold_df[FINDING_COLUMNS] = gold_df[FINDING_COLUMNS].astype('Int64')\n\ngold_train_ids, gold_test_ids = train_test_split(gold_df[['StudyInstanceUID', 'Is_English']], test_size = 0.33, stratify = gold_df['Is_English'], random_state = 42)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:05.215066Z","iopub.execute_input":"2026-09-17T20:14:05.215443Z","iopub.status.idle":"2026-09-17T20:14:05.452704Z","shell.execute_reply.started":"2026-09-17T20:14:05.215404Z","shell.execute_reply":"2026-09-17T20:14:05.451608Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"support_by_lang_df = (gold_df\n    .assign(set = lambda x: np.where(x['StudyInstanceUID'].isin(gold_train_ids['StudyInstanceUID']), 'train', 'test'))\n    .groupby(['set','Is_English'])[FINDING_COLUMNS]\n    .sum()\n    .reset_index())\nfinding_order_by_support = support_by_lang_df[FINDING_COLUMNS].sum().sort_values(ascending=False).index.tolist()\nsupport_by_lang_long_df = (support_by_lang_df\n    .melt(id_vars = ['set', 'Is_English'], var_name=\"Classification\", value_name=\"Count\", ignore_index=False)  \n)\n\n(\n    ggplot(support_by_lang_long_df, aes(x=\"Classification\", y=\"Count\", fill = 'set'))\n        + geom_col(position=\"dodge\", color=\"white\")\n        + facet_wrap('Is_English')\n        + scale_x_discrete(limits=finding_order_by_support)\n        + scale_fill_manual(\n            values={'train': '#1f77b4', 'test': '#878f99',},\n            name=\"Set\",\n        )\n        + labs(x=\"Class\", y=\"Support, n\", title=\"Support by Class (All Labeled Data)\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1), figure_size=(10, 6))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:05.453835Z","iopub.execute_input":"2026-09-17T20:14:05.454163Z","iopub.status.idle":"2026-09-17T20:14:06.268611Z","shell.execute_reply.started":"2026-09-17T20:14:05.454134Z","shell.execute_reply":"2026-09-17T20:14:06.267606Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Just a basic reminder that no matter what the outcome there is not a *lot* of training\n  data and the classes are pretty imbalanced.\n- For an initial look, it will make sense to split on English vs Non-English, but once\n  we start testing for optimization we probably need to be able to look at label performance\n  without respect to language, or we'd need to focus on on a subset of classes.","metadata":{}},{"cell_type":"markdown","source":"## 2. Baseline LLM Labeling Performance - All Labeled Data\n\n- Generally, how well do LLMs recover the gold labels at baseline (eg, simple prompt, no optimization)?  For this first pass we'll also look at English vs Non-English reports separately\n  to understand how language differences impact model performance.  \n- For this first test, the instructions will be fairly naive - just look for indications of injury/presence of the label in the report, which will provide insight into how each model\n  can understand the clinical context without respect to strict criteria used for independent labeleing that may or may not be present report text\n","metadata":{}},{"cell_type":"code","source":"def get_scores_by_class(gold_df: pd.DataFrame, pred_df: pd.DataFrame, findings: list = FINDING_COLUMNS, na_conversion: int | None = None) -> dict:\n    golds = gold_df.copy()\n    if na_conversion is not None:\n        preds = pred_df.copy().fillna(na_conversion)\n        print(f\"{pred_df[findings].isna().to_numpy().sum()} missing labels removed - now {preds[findings].isna().to_numpy().sum()} missing labels.\")\n    else:\n        preds = pred_df.copy()\n    scores_dict = {}\n    for col in findings: \n        pred = preds[['StudyInstanceUID', col]].dropna().sort_values('StudyInstanceUID')\n        uids = pred['StudyInstanceUID'].to_list()\n        gold = golds[golds['StudyInstanceUID'].isin(uids)].sort_values('StudyInstanceUID')[col]\n        scores_dict[col] = {\n            'precision': precision_score(gold, pred[col]),\n            'recall': recall_score(gold, pred[col]),\n            'specificity': recall_score(gold, pred[col], pos_label = 0),\n            'f1-score': f1_score(gold, pred[col]),\n            'npv': precision_score(gold, pred[col], pos_label = 0),\n            'balanced_accuracy_score': balanced_accuracy_score(gold, pred[col]),\n            'support': gold[gold == 1].shape[0],\n            'total_labeled': gold.shape[0]\n        }\n    return scores_dict\n\ndef get_confusion_matrices_by_class(gold_df: pd.DataFrame, pred_df: pd.DataFrame, findings: list = FINDING_COLUMNS, na_conversion: int | None = None) -> dict:\n    golds = gold_df.copy()\n    if na_conversion is not None:\n        preds = pred_df.copy().fillna(na_conversion)\n        print(f\"{pred_df[findings].isna().to_numpy().sum()} missing labels removed - now {preds[findings].isna().to_numpy().sum()} missing labels.\")\n    else:\n        preds = pred_df.copy()\n    mtx_dict = {}\n    for col in findings:  \n        pred = preds[['StudyInstanceUID', col]].dropna().sort_values('StudyInstanceUID')\n        uids = pred['StudyInstanceUID'].to_list()\n        gold = golds[golds['StudyInstanceUID'].isin(uids)].sort_values('StudyInstanceUID')[col]\n        mtx_dict[col] = confusion_matrix(gold, pred[col], labels = [0,1])\n    return mtx_dict\n\ndef make_cohens_kappa_matrix_plot(preds_long_df: pd.DataFrame, models: dict[str, ModelSpec] = MODELS, model_col: str = 'Model', annotation_col:str = 'annotation', scale_limits: tuple[float] = (0.5, 0.95), mid_value:float | None = None) -> ggplot:\n    model_keys = list(models)\n    agreements = {}\n    for i, model_a in enumerate(model_keys):\n        agreement = {}\n        for model_b in model_keys[i+1:]:\n            agreement[model_b] = cohen_kappa_score(\n                y1=preds_long_df.loc[\n                    preds_long_df[model_col] == model_a, annotation_col],\n                y2=preds_long_df.loc[\n                    preds_long_df[model_col] == model_b, annotation_col],\n                labels=[-1, 0, 1],\n            )\n        agreements[model_a] = agreement\n\n    kappa_df = pd.DataFrame(agreements).T.reindex(index=model_keys, columns=model_keys)\n    kappa_df = kappa_df.combine_first(kappa_df.T)                       \n    kappa_df = kappa_df.mask(np.eye(len(kappa_df), dtype=bool), 1.0)    \n\n    kappa_long_df = (kappa_df.rename_axis('Model_A')\n        .reset_index()\n        .melt(id_vars='Model_A', var_name='Model_B', value_name='kappa')\n        .dropna(subset=['kappa']))\n\n    off_diag = kappa_long_df.loc[kappa_long_df['Model_A'] != kappa_long_df['Model_B'], 'kappa']\n    \n    mid = mid_value if mid_value is not None else off_diag.median()\n       \n\n    kappa_long_df['Model_A'] = pd.Categorical(kappa_long_df['Model_A'], categories=model_keys, ordered=True)\n    kappa_long_df['Model_B'] = pd.Categorical(kappa_long_df['Model_B'], categories=model_keys[::-1], ordered=True)\n\n    return (\n                ggplot(kappa_long_df, aes('Model_A', 'Model_B', fill='kappa'))\n                    + geom_tile(color='white', size=0.5)\n                    + geom_text(aes(label='kappa'), format_string='{:.2f}', size=8)\n                    + scale_fill_gradient2(low='#4575b4', mid='#ffffbf', high='#d73027',\n                                    midpoint=mid, limits= scale_limits)\n                    + coord_fixed()\n                    + labs(x='', y='', fill=\"Cohen's kappa\", title = 'Inter-Rater (LLM) Reliability Matrix')\n                    + theme_minimal()\n                    + theme(axis_text_x=element_text(angle=45, hjust=1))\n            )\n\n\ndef map_annotations_to_probabilities(gold_df: pd.DataFrame, pred_df: pd.DataFrame, scores_df: pd.DataFrame, findings: list = FINDING_COLUMNS, abstain_probability: float = 0.5) -> pd.DataFrame:\n    \"\"\"Score each hard annotation by how reliable that call is for its class - a positive call takes\n    the class PPV, a negative call takes 1 - NPV, and an abstention takes `abstain_probability`.\n    `scores_df` supplies `precision` and `npv` per Model and Finding.\"\"\"\n    preds_long_df = (pred_df[['StudyInstanceUID', 'Model'] + findings]\n        .melt(id_vars = ['StudyInstanceUID', 'Model'], var_name = 'Finding', value_name = 'annotation'))\n    golds_long_df = (gold_df[['StudyInstanceUID'] + findings]\n        .melt(id_vars = 'StudyInstanceUID', var_name = 'Finding', value_name = 'gold_label'))\n\n    probabilities_df = (preds_long_df\n        .merge(golds_long_df, on = ['StudyInstanceUID', 'Finding'], how = 'left')\n        .merge(scores_df[['Model', 'Finding', 'precision', 'npv']], on = ['Model', 'Finding'], how = 'left'))\n    probabilities_df['annotation'] = probabilities_df['annotation'].astype('Int64')\n    is_positive = probabilities_df['annotation'].eq(1).fillna(False).to_numpy(dtype = bool)\n    is_negative = probabilities_df['annotation'].eq(0).fillna(False).to_numpy(dtype = bool)\n    probabilities_df['probability'] = np.where(\n        is_positive,\n        probabilities_df['precision'],\n        np.where(is_negative, 1 - probabilities_df['npv'], abstain_probability)\n    )\n    return probabilities_df[['StudyInstanceUID', 'Model', 'Finding', 'gold_label', 'annotation', 'probability']]\n\n# This is basic setup to use dspy - define the input classes -> to get the dspy signature\n# dspy signature -> to get the dspy module; Example formatting fuction to shape the input\n# data as dspy expects\n\ndef format_examples_baseline(df: pd.DataFrame) -> list[dspy.Example]:\n    df_findings = df[FINDING_COLUMNS].astype(int).astype(bool).to_dict(orient = 'records')\n    examples = []\n    for i, row in enumerate(df.itertuples()):\n        example = dspy.Example(\n            study_instance_uid = row.StudyInstanceUID,\n            report = row.Report,\n            findings = df_findings[i]\n        ).with_inputs(\"report\", \"study_instance_uid\")\n        examples.append(example)\n    return examples\n\nclass ReportFindingsBaseline(TypedDict):\n    ACL: bool\n    MCL: bool\n    Medial_Meniscus: bool\n    Lateral_Meniscus: bool\n    Medial_OA: bool\n    Lateral_OA: bool\n    PF_OA: bool\n    Effusion: bool\n    Synovitis: bool\n    Baker: bool\n    Contusion: bool\n    Fracture: bool\n\nclass KneeFindingsBaseline(BaseModel):\n    findings: ReportFindingsBaseline = Field(description = (\n\"ACL: anterior cruciate ligament.\\n\"\n\"MCL: medial collateral ligament.\\n\"\n\"Medial_Meniscus: medial meniscus.\\n\"\n\"Lateral_Meniscus: lateral meniscus.\\n\"\n\"Medial_OA: medial tibiofemoral compartment osteoarthritis.\\n\"\n\"Lateral_OA: lateral tibiofemoral compartment osteoarthritis.\\n\"\n\"PF_OA: patellofemoral compartment osteoarthritis.\\n\"\n\"Effusion: joint effusion.\\n\"\n\"Synovitis: synovitis.\\n\"\n\"Baker: Baker's (popliteal) cyst.\\n\"\n\"Contusion: bone contusion.\\n\"\n\"Fracture: fracture.\"\n        )\n    )\n    rationale: str = Field(description = (\n\"Reasoning for each of the classifcations based on the\"\n\"radiology report contents\"\n        )\n    )\n\n# In dspy, the docstring of a signature is used as the 'instructions' component of a prompt\n# Attempt here is to be basic, and leave open the potential for dspy optimization down road\nclass KneeMRIOutputBaseline(dspy.Signature):\n    \"\"\"\n    Classify a knee MRI radiology report for the presence of:\n    - findings of injury to the anterior cruciate ligament (ACL), \n      medial collateral ligament (MCL), medial meniscus, and/or lateral meniscus\n    - findings of medial tibiofemoral compartment osteoarthritis (Medial OA), \n      patellofemoral compartment osteoarthritis (PF OA), effusion, synovitis, Bakers' (popliteal)\n      Cyst, contusion, and/or fracture\n\n    Reports may be in any language and may follow any template. Read the entire\n    report carefully before drawing conclusions.\n    \"\"\"\n    report: str = dspy.InputField()\n    classification: KneeFindingsBaseline = dspy.OutputField()\n\nclass BaselineClassifier(dspy.Module):\n    def __init__(self):\n        super().__init__()\n        self.classification_module = dspy.ChainOfThought(KneeMRIOutputBaseline)\n        \n    def forward(self, report: str, study_instance_uid: str) -> dspy.Prediction: \n        result = self.classification_module(report=report) \n        \n        return dspy.Prediction(\n            study_instance_uid = study_instance_uid,\n            findings = result.classification.findings, \n            rationale = result.classification.rationale\n        )\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:06.269946Z","iopub.execute_input":"2026-09-17T20:14:06.27039Z","iopub.status.idle":"2026-09-17T20:14:06.310103Z","shell.execute_reply.started":"2026-09-17T20:14:06.270341Z","shell.execute_reply":"2026-09-17T20:14:06.309047Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Runs predictions using the baseline prompt acrosst the 8 models, unless the serialized\n# output is available\n\nbaseline_preds_load_path = DATA_LOC / 'datasets' / 'joshuaziel' / 'rsna-knee-mri-radiology-report-classifications' / 'raw_prediction_results.json'\nbaseline_preds_persist_path = OUTPUTS_LOC / 'raw_prediction_results.json'\n\nif baseline_preds_load_path.is_file():\n    with open(str(baseline_preds_load_path), 'r') as f:\n        baseline_preds_raw = json.load(f)\n    print(\"Loaded LM predictions from file\")\n\n# Note that there is a batch processing path in dspy, however, Anthropic models in our experience have been flaky\n# with the litellm backend in dspy.  Instead we'll just hand-roll a basic retry loop.\nelse:        \n    program = BaselineClassifier()\n    baseline_preds_raw = {}\n    for model_key, model_spec in tqdm(MODELS.items(), dynamic_ncols = True):\n        dspy_lm = dspy.LM(model = model_spec.name, temperature = 1.0, model_type = model_spec.type, api_key = model_spec.key, cache = False) \n        #  # Note: This will now break for Fable for some reason, unless tool_choice = 'auto' is added, which will break it for OpenAI\n        dspy.configure(lm=dspy_lm)\n        model_preds = []\n        \n        for row in gold_df.itertuples():\n            success, tries = False, 0\n            while not success and tries < 4:\n                try:\n                    prediction = program(study_instance_uid = row.StudyInstanceUID, report = row.Report)\n                    model_preds.append({'StudyInstanceUID': prediction.study_instance_uid, **prediction.findings, 'rationale': prediction.rationale})\n                    success = True\n                    tries += 1\n                except Exception as e:\n                    tqdm.write(f'Error predicting labels for {row.StudyInstanceUID}: {e}')\n                    tries += 1\n            if not success:\n                missing_dict = {key: None for key in FINDING_COLUMNS}\n                model_preds.append({'StudyInstanceUID': row.StudyInstanceUID, **missing_dict, 'rationale': None}) \n        baseline_preds_raw[model_key] = model_preds\n    \n    with open(str(baseline_preds_persist_path), 'w') as f:\n        json.dump(baseline_preds_raw, f, indent = 4)\n    \n# Get the data into dataframe\nbaseline_preds_parts = []\nfor model_key, model_records in baseline_preds_raw.items():\n    for record in model_records:\n        record['Model'] = model_key\n        baseline_preds_parts.append(record)\nbaseline_preds_wide_df = pd.DataFrame(baseline_preds_parts).merge(gold_df[['StudyInstanceUID', 'Is_English']], on = 'StudyInstanceUID', how = \"left\")\nbaseline_preds_wide_df = baseline_preds_wide_df[baseline_preds_wide_df['StudyInstanceUID'].isin(gold_train_ids['StudyInstanceUID'])]\nbaseline_preds_wide_df[FINDING_COLUMNS] = baseline_preds_wide_df[FINDING_COLUMNS].astype(int)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:06.312759Z","iopub.execute_input":"2026-09-17T20:14:06.31315Z","iopub.status.idle":"2026-09-17T20:14:06.363955Z","shell.execute_reply.started":"2026-09-17T20:14:06.313101Z","shell.execute_reply":"2026-09-17T20:14:06.362844Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"make_cohens_kappa_matrix_plot(\n    baseline_preds_wide_df.melt(\n        id_vars = ['StudyInstanceUID', 'Model', 'rationale', 'Is_English'],\n        var_name = 'Finding',\n        value_name = 'annotation'\n    ), \n    model_col = 'Model', \n    annotation_col = 'annotation',\n    mid_value = 0.7, # 0.7 might generally be consider to be a moderate level of agreement\n    scale_limits = (0.00, 1.0)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:06.36515Z","iopub.execute_input":"2026-09-17T20:14:06.365503Z","iopub.status.idle":"2026-09-17T20:14:07.172374Z","shell.execute_reply.started":"2026-09-17T20:14:06.365472Z","shell.execute_reply":"2026-09-17T20:14:07.171351Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Wow, i did not expect the LLMs to be so similar across the board!  This probably isnt a 'hard' task for them generally,\n  given that the least and most capable models from both Anthropic and OpenAI all agree pretty substantially. \n","metadata":{}},{"cell_type":"code","source":"baseline_bylang_score_parts = []\nfor model_key in MODEL_KEYS:\n    for lang_group in LANG_GROUPS:\n        model_scores_df = pd.DataFrame(get_scores_by_class(\n            gold_df,\n            baseline_preds_wide_df[(baseline_preds_wide_df['Model'] == model_key) & (baseline_preds_wide_df['Is_English'] == lang_group)]\n        )).rename_axis('metric').reset_index(level = 'metric')\n        model_scores_df['Model'] = model_key\n        model_scores_df['Is_English'] = lang_group\n        baseline_bylang_score_parts.append(model_scores_df)\nbaseline_bylang_scores_df = (\n    pd.concat(baseline_bylang_score_parts)\n        .melt(id_vars = ['Model', 'metric', 'Is_English'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['Model', 'Finding', 'Is_English'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\nbaseline_bylang_scores_df[['support', 'total_labeled']] = baseline_bylang_scores_df[['support', 'total_labeled']].astype(\"Int64\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:07.173578Z","iopub.execute_input":"2026-09-17T20:14:07.173885Z","iopub.status.idle":"2026-09-17T20:14:11.636899Z","shell.execute_reply.started":"2026-09-17T20:14:07.173855Z","shell.execute_reply":"2026-09-17T20:14:11.635773Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(\n        (\n            baseline_bylang_scores_df[['Model', 'Is_English', 'f1-score', 'precision', 'recall', 'balanced_accuracy_score', 'npv', 'specificity']]\n                .groupby(['Model', 'Is_English'])\n                .mean()\n                .reset_index()\n                .melt(id_vars = ['Model', 'Is_English'], var_name = 'Metric', value_name = 'Value')\n        ),\n        aes(x = 'Model', y = 'Value', fill = 'Is_English')\n    )\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Metric', ncol = 3)\n        + scale_x_discrete(limits = MODEL_KEYS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'English': '#1f77b4', 'Non-English/Unknown': '#878f99',},\n            name=\"Report Language\",\n        )\n        + labs(x=\"LLM Model\", y=\"Macro Average\", title=\"Overall Performance by Model and Language (English vs Non-English)\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(10, 5))\n\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:11.638055Z","iopub.execute_input":"2026-09-17T20:14:11.638396Z","iopub.status.idle":"2026-09-17T20:14:12.577352Z","shell.execute_reply.started":"2026-09-17T20:14:11.638364Z","shell.execute_reply":"2026-09-17T20:14:12.576319Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- As expected based on the agreement matrix, macro average performance is highly similar across models; whether the report\n  text is in English appears to matter little overall","metadata":{}},{"cell_type":"code","source":"(\n    ggplot(baseline_bylang_scores_df, aes(x = 'Model', y = 'precision', fill = 'Is_English'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = MODEL_KEYS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'English': '#1f77b4', 'Non-English/Unknown': '#878f99',},\n            name=\"Report Language\",\n        )\n        + labs(x=\"LLM Model\", y=\"Precision\", title=\"Positive Predictive Value (Precision) by Model and Language (English vs Non-English)\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:12.578622Z","iopub.execute_input":"2026-09-17T20:14:12.5801Z","iopub.status.idle":"2026-09-17T20:14:14.47719Z","shell.execute_reply.started":"2026-09-17T20:14:12.580052Z","shell.execute_reply":"2026-09-17T20:14:14.476102Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(baseline_bylang_scores_df, aes(x = 'Model', y = 'npv', fill = 'Is_English'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = MODEL_KEYS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'English': '#1f77b4', 'Non-English/Unknown': '#878f99',},\n            name=\"Report Language\",\n        )\n        + labs(x=\"LLM Model\", y=\"NPV\", title=\"Negative Predictive Value (NPV) by Model and Language (English vs Non-English)\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:14.478411Z","iopub.execute_input":"2026-09-17T20:14:14.47872Z","iopub.status.idle":"2026-09-17T20:14:16.043048Z","shell.execute_reply.started":"2026-09-17T20:14:14.47869Z","shell.execute_reply":"2026-09-17T20:14:16.041747Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n  - The fact that there seems to be such a similar ceiling on many classes is very consistent with the overall level of agreement between models;\n    my guess is that most gains - if any - are going to come from applying specific criteria, rather than prompt hacks to make model reasoning stronger\n    but that remains to be seen.\n  - Larger operative conclusion - there is essentially no consistent data suggesting that any one of the general-purpose LLMs performs meaningfully\n    better or worse on the baseline condition\n  - Language (English vs Non-English) similarly did not seem to matter consistently in terms of performance; in fact in some cases, the Non-English subset was \n    categorized numerically better.  Given the likely size of the confidence intervals here, this is essentially noise.\n\n## Thoughts on Individual Classes\n- Recall is pretty good across the board essentially perfect across the board for ACL, Effusion, MCL, Lateral & Medial Meniscus (English), Lateral OA (Non-English), and Medial OA  \n  (Non English); Generally, with the clear exception of Synovitis (which is a big one) the problem across the board is precision. \n- In some cases its pretty bad: PPV for Lateral OA and MCL are basically guess territory; contusion and effusion aren't much better\n\n","metadata":{}},{"cell_type":"markdown","source":"## 3. Repeat LLM Classification of Labeled Data Based on Reports\n- Provide the competition labeling definitions for the classes explicitly as part of the output schema\n- Update program to allow for an 'Undetermined' Class\n- Capture a confidence self-assessment and rationale for each class on each report\n- No real training is occurring, so this will be done on all labeled train data","metadata":{}},{"cell_type":"code","source":"class ClassificationLabel(Enum):\n    PRESENT = 1\n    ABSENT = 0\n    UNDETERMINED = None\n\n\nclass ConfidenceRating(Enum):\n    VERY_HIGH = 5\n    HIGH = 4\n    MODERATE = 3\n    LOW = 2\n    VERY_LOW = 1\n    NO_CONFIDENCE = 0\n\n\nclass FindingAnnotation(TypedDict):\n    annotation: Annotated[\n        ClassificationLabel,\n        Field(\n            description=\n\"\"\"PRESENT, ABSENT, or UNDETERMINED for this finding, per the finding definition.\"\"\"\n        ),\n    ]\n    confidence: Annotated[\n        ConfidenceRating,\n        Field(\n            description=\n\"\"\"Confidence in the assigned annotation, based on the explicitness and internal consistency\nof the report text.\"\"\"\n        ),\n    ]\n    rationale: Annotated[\n        str,\n        Field(\n            description=\n\"\"\"Short justification (in English, regardless of the report language) for the assigned\nannotation and confidence.\"\"\"\n        ),\n    ]\n\n\nclass ReportFindingsCriteria(TypedDict):\n    ACL: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A high-grade partial or full-thickness tear of the anterior cruciate ligament, meaning complete discontinuity\nof the ligament, or more than 50 percent of fibers disrupted, with or without secondary signs such as characteristic\npivot-shift bone contusions. Mild signal change, degeneration, or thickening without discontinuity is graded negative.\"\"\"\n        ),\n    ]\n    MCL: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A high-grade partial or complete acute tear of the medial collateral ligament, with disrupted fibers\nand edema within and adjacent to the ligament. Low-grade sprains and chronic or remote stress changes are graded negative.\"\"\"\n        ),\n    ]\n    Medial_Meniscus: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the medial meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Lateral_Meniscus: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the lateral meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Medial_OA: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the medial compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Lateral_OA: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the lateral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    PF_OA: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the patellofemoral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Effusion: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large amount of fluid distending the joint.\"\"\"\n        ),\n    ]\n    Synovitis: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"Inflammation and thickening of the synovial lining of the joint.\"\"\"\n        ),\n    ]\n    Baker: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large fluid collection in the characteristic Baker (popliteal) cyst location behind the knee.\"\"\"\n        ),\n    ]\n    Contusion: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A bone contusion, seen as bone marrow edema-like signal from impact, without a discrete fracture line.\"\"\"\n        ),\n    ]\n    Fracture: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"An acute cortical break or fracture line.\"\"\"\n        ),\n    ]\n\n\n# Same 12 names as FINDING_COLUMNS, derived from the schema so the two cannot drift apart\nFINDING_NAMES = tuple(ReportFindingsCriteria.__annotations__)\nENUM_KEYS = {'annotation': ClassificationLabel, 'confidence': ConfidenceRating}\n\n\ndef _as_int(value, enum_cls: type[Enum]) -> int | None:\n    if value is None:\n        return None\n    if isinstance(value, enum_cls):\n        return None if value.value is None else int(value.value)\n    if isinstance(value, str):\n        name = value.strip().upper()\n        if name in enum_cls.__members__:\n            member = enum_cls[name]\n            return None if member.value is None else int(member.value)\n        value = int(value)\n    member = enum_cls(int(value))\n    return None if member.value is None else int(member.value)\n\n\ndef as_long_records(findings: dict, study_instance_uid: str) -> list[dict]:\n    records = []\n    for name in FINDING_NAMES:\n        annotation = findings.get(name) or {}\n        rationale = annotation.get('rationale')\n        records.append({\n            'StudyInstanceUID': str(study_instance_uid),\n            'Finding': name,\n            **{key: _as_int(annotation.get(key), cls) for key, cls in ENUM_KEYS.items()},\n            'rationale': None if rationale is None else str(rationale),\n        })\n    return records\n\n\nclass KneeFindingsCriteria(BaseModel):\n    findings: ReportFindingsCriteria = Field(\n        description=\n\"\"\"Annotations per class for the knee MRI radiology report\"\"\"\n    )\n\n\nclass KneeMRIOutputCriteria(dspy.Signature):\n    \"\"\"\n    - Act as an expert musculoskeletal radiologist to annotate a knee MRI radiology report for each class of finding:\n      - Use the definitions for each class of finding when making annotations\n      - **Classes of findings**: anterior cruciate ligament (ACL) tear, medial collateral ligament (MCL)\n        tear, medial meniscus tear, lateral meniscus tear, medial tibiofemoral compartment osteoarthritis (Medial_OA),\n        lateral tibiofemoral compartment osteoarthritis (Lateral_OA), patellofemoral compartment osteoarthritis (PF_OA),\n        effusion, synovitis, Bakers' (popliteal) cyst, contusion, and/or fracture\n      - *Annotations for each class*:\n        - **PRESENT**: Based on the text of the radiology report and the class definition, the finding is clearly identified or it can be\n          reasonably inferred if not named directly.\n        - **ABSENT**: Based on the text of the radiology report and the class definition, the finding is clearly excluded or its exclusion can be\n          reasonably inferred if not named directly.\n        - **UNDETERMINED**: Based on the text of the radiology report and the class definition, a determination of **PRESENT** or **ABSENT**\n          cannot be made.\n      - Each class takes an object with an annotation, a confidence rating, and a short rationale citing the report language used.\n    \"\"\"\n    report: str = dspy.InputField()\n    classification: KneeFindingsCriteria = dspy.OutputField()\n\n\nclass CriteriaClassifier(dspy.Module):\n    def __init__(self):\n        super().__init__()\n        self.classification_module = dspy.ChainOfThought(KneeMRIOutputCriteria)\n\n    def forward(self, report: str, study_instance_uid: str) -> dspy.Prediction:\n        result = self.classification_module(report=report)\n\n        return dspy.Prediction(\n            study_instance_uid=study_instance_uid,\n            findings=as_long_records(result.classification.findings, study_instance_uid),\n        )\n\ndef format_examples_criteria(df: pd.DataFrame) -> list[dspy.Example]:\n    df_findings = df[FINDING_COLUMNS].astype(int).to_dict(orient = 'records')\n    examples = []\n    for i, row in enumerate(df.itertuples()):\n        example = dspy.Example(\n            study_instance_uid = row.StudyInstanceUID,\n            report = row.Report,\n            findings = df_findings[i]\n        ).with_inputs(\"report\", \"study_instance_uid\")\n        examples.append(example)\n    return examples\n# Runs predictions using the criteria prompt acrosst the 8 models, unless the serialized\n# output is available\n\ncriteria_preds_load_path = DATA_LOC / 'datasets' / 'joshuaziel' / 'rsna-knee-mri-radiology-report-classifications' / 'raw_prediction_results_v2.json'\ncriteria_preds_persist_path = OUTPUTS_LOC / 'raw_prediction_results_v2.json'\n\nif criteria_preds_load_path.is_file():\n    with open(str(criteria_preds_load_path), 'r') as f:\n        criteria_preds_raw = json.load(f)\n    print(\"Loaded LM predictions from file\")\n\n# Note that as before there is a batch processing path in dspy, however, Anthropic models in our experience have been flaky\n# with the litellm backend in dspy.  Instead we're using the hand-rolled retry loop.\nelse:        \n    program = CriteriaClassifier()\n    criteria_preds_raw = {}\n    for model_key, model_spec in tqdm(MODELS.items(), dynamic_ncols = True):\n        dspy_lm = dspy.LM(model = model_spec.name, temperature = 1.0, model_type = model_spec.type, api_key = model_spec.key, cache = False)\n         # Note: This will now break for Fable for some reason, unless tool_choice = 'auto' is added.\n        dspy.configure(lm=dspy_lm)\n        model_preds = []\n\n        for row in gold_df.itertuples():\n            success, tries = False, 0\n            while not success and tries < 4:\n                tries += 1\n                try:\n                    prediction = program(study_instance_uid=row.StudyInstanceUID, report=row.Report)\n                    model_preds.extend(prediction.findings)\n                    success = True\n                except Exception as e:\n                    tqdm.write(f'Error predicting labels for {row.StudyInstanceUID}: {e}')\n            if not success:\n                model_preds.extend(as_long_records({}, row.StudyInstanceUID))\n        criteria_preds_raw[model_key] = model_preds\n    \n    with open(str(criteria_preds_persist_path), 'w') as f:\n        json.dump(criteria_preds_raw, f, indent = 4)\n\ncriteria_preds_long_df = pd.concat(\n    [pd.DataFrame(records).assign(Model=model_key) for model_key, records in criteria_preds_raw.items()],\n    ignore_index=True,\n)\ncriteria_preds_long_df = criteria_preds_long_df[criteria_preds_long_df['StudyInstanceUID'].isin(gold_train_ids['StudyInstanceUID'])]\ncriteria_preds_long_df['annotation'] = criteria_preds_long_df['annotation'].astype(\"Int64\")\ncriteria_preds_long_df['Undetermined'] = pd.isna(criteria_preds_long_df['annotation'])\ncriteria_preds_long_filled_df = criteria_preds_long_df.copy()\ncriteria_preds_long_filled_df.loc[criteria_preds_long_filled_df['Undetermined'], 'annotation'] = -1\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:16.044303Z","iopub.execute_input":"2026-09-17T20:14:16.044648Z","iopub.status.idle":"2026-09-17T20:14:16.142835Z","shell.execute_reply.started":"2026-09-17T20:14:16.044606Z","shell.execute_reply":"2026-09-17T20:14:16.141811Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 4. Re-Evaluation of LLM Agreement, Confidence, and Undetermined Labels (Including Relative to Baseline)\n- How well do the LLMs agree and are there any patterns by provider?\n- How confident are the LLMs by label, and are there any patterns?\n- Which LLMs provide the most labels by class, which the least?\n- How does performance change relative to baseline?","metadata":{}},{"cell_type":"code","source":"make_cohens_kappa_matrix_plot(criteria_preds_long_filled_df, mid_value = 0.7, scale_limits = (0.00, 1.0)) # 0.7 might generally be consider to be a moderate level of agreement","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:16.143957Z","iopub.execute_input":"2026-09-17T20:14:16.144301Z","iopub.status.idle":"2026-09-17T20:14:17.209063Z","shell.execute_reply.started":"2026-09-17T20:14:16.144258Z","shell.execute_reply":"2026-09-17T20:14:17.20804Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Substantially more disagreement here, which means there could be more signal on performance metrics.  Nontheless, agreement is generally still good among models from Anthropic and OpenAI, respectively \n  - Within models from each provider, agreement is generally Strong, or Moderate (see [McHugh, 2012](https://pmc.ncbi.nlm.nih.gov/articles/PMC3900052/))\n- Highest agreement overall between Opus and Fable, with reasonably similar results for Luna and Sol or Terra and Sol from OpenAI\n- Nano is a clear outlier, demonstrating only Moderate (or Weak) agreement with all other models","metadata":{}},{"cell_type":"code","source":"(\n    ggplot(criteria_preds_long_df, aes(x=\"confidence\"))\n        + geom_histogram(bins = 6, color = \"white\", fill = \"#1f77b4\")\n        + facet_grid(cols = 'Finding', rows = 'Model')\n        + scale_x_continuous(breaks=[1,3,5], labels=[\"Very Low\", \"Moderate\", \"Very High\"])\n        + theme(figure_size=(15, 10))\n        + labs(x=\"Confidence Level\", y=\"Count, n\", title=\"Frequency Distribution of Confidence Levels by LLM and Class\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:17.210145Z","iopub.execute_input":"2026-09-17T20:14:17.210484Z","iopub.status.idle":"2026-09-17T20:14:25.538112Z","shell.execute_reply.started":"2026-09-17T20:14:17.210452Z","shell.execute_reply":"2026-09-17T20:14:25.53707Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Generally models appear to provide a range of self-reported confidence levels without clearly collapsing to a mean value\n  - Visually, nano appears to provide the widest confidence spreads across classes, haiku the least\n- Models self-report the most confidence on ACL, MCL, and Medial_Meniscus across the board (which does include confidence on 'undetermined')","metadata":{}},{"cell_type":"code","source":"model_order_by_abstention = criteria_preds_long_df[['Model', 'Undetermined']].groupby('Model').sum().sort_values(by = 'Undetermined', ascending = True).index.to_list()\n(\n    ggplot(\n        criteria_preds_long_df.groupby(['Model', 'Finding'])['Undetermined'].sum().reset_index(),\n        aes(x = 'Model', y = 'Undetermined')\n    )\n    + facet_wrap('Finding')\n    + geom_col(position=\"dodge\", color=\"white\", fill = \"#1f77b4\")\n    + scale_x_discrete(limits=model_order_by_abstention)\n    + theme(figure_size=(8, 6))\n    + labs(x=\"LLM Model\", y=\"Labeling Abstentions, n\", title=\"Frequency of Labeling Abstention by Class and LLM\")\n    + theme(axis_text_x=element_text(angle=45, hjust=1))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:25.541964Z","iopub.execute_input":"2026-09-17T20:14:25.542336Z","iopub.status.idle":"2026-09-17T20:14:26.708251Z","shell.execute_reply.started":"2026-09-17T20:14:25.542303Z","shell.execute_reply":"2026-09-17T20:14:26.707189Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Broadly speaking, models from OpenAI decline to label the largest numbers of classes; Anthropic and OpenAI models\n  appear to reasoning to different conclusions about what to abstain from\n- On specific classes, LLMs appear to decline labeling Baker Cyst, Medial or Lateral OA, and Synovitis more frequently\n  than other classes\n- Declining could be good or bad - need to evaluate the impact on performance metrics","metadata":{}},{"cell_type":"code","source":"criteria_preds_wide_df = criteria_preds_long_df.pivot(index = ['Model', 'StudyInstanceUID'], columns = 'Finding', values = 'annotation').reset_index()\n\nbaseline_score_parts = []\nfor model_key in MODEL_KEYS:\n    model_scores_df = pd.DataFrame(get_scores_by_class(\n        gold_df,\n        baseline_preds_wide_df[baseline_preds_wide_df['Model'] == model_key]\n    )).rename_axis('metric').reset_index(level = 'metric')\n    model_scores_df['Model'] = model_key\n    baseline_score_parts.append(model_scores_df)\nbaseline_scores_df = (\n    pd.concat(baseline_score_parts)\n        .melt(id_vars = ['Model', 'metric'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['Model', 'Finding'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\nbaseline_scores_df['support_pct'] = baseline_scores_df['support'] / baseline_scores_df['support'] \nbaseline_scores_df['Condition'] = 'baseline'\n\ncriteria_filled_score_parts = []\nfor model_key in MODEL_KEYS:\n    model_scores_df = pd.DataFrame(get_scores_by_class(\n        gold_df,\n        criteria_preds_wide_df[criteria_preds_wide_df['Model'] == model_key],\n        na_conversion = 0\n    )).rename_axis('metric').reset_index(level = 'metric')\n    model_scores_df['Model'] = model_key\n    criteria_filled_score_parts.append(model_scores_df)\ncriteria_filled_scores_df = (\n    pd.concat(criteria_filled_score_parts)\n        .melt(id_vars = ['Model', 'metric'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['Model', 'Finding'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\ncriteria_filled_scores_df['support_pct'] = criteria_filled_scores_df['support'] / baseline_scores_df['support'] \ncriteria_filled_scores_df['Condition'] = 'all-filled'\n\ncriteria_masked_score_parts = []\nfor model_key in MODEL_KEYS:\n    model_scores_df = pd.DataFrame(get_scores_by_class(\n        gold_df,\n        criteria_preds_wide_df[criteria_preds_wide_df['Model'] == model_key]\n    )).rename_axis('metric').reset_index(level = 'metric')\n    model_scores_df['Model'] = model_key\n    criteria_masked_score_parts.append(model_scores_df)\ncriteria_masked_scores_df = (\n    pd.concat(criteria_masked_score_parts)\n        .melt(id_vars = ['Model', 'metric'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['Model', 'Finding'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\ncriteria_masked_scores_df['support_pct'] = criteria_masked_scores_df['support'] / baseline_scores_df['support'] \ncriteria_masked_scores_df['Condition'] = 'only-labeled'\n\ncriteria_scores_df = pd.concat([baseline_scores_df, criteria_filled_scores_df, criteria_masked_scores_df])\ncriteria_scores_df[['support', 'total_labeled']] = criteria_scores_df[['support', 'total_labeled']].astype('Int64')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:26.709479Z","iopub.execute_input":"2026-09-17T20:14:26.709775Z","iopub.status.idle":"2026-09-17T20:14:34.092524Z","shell.execute_reply.started":"2026-09-17T20:14:26.709737Z","shell.execute_reply":"2026-09-17T20:14:34.091516Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 6. Prediction Performance Overall and by Class Across Models\n- How do the updated instructions impact overall performance by Model?\n- What does PPV and NPV look like at the class level, given that we will need these for weak supervision?","metadata":{}},{"cell_type":"code","source":"(\n    ggplot((criteria_scores_df\n            .loc[criteria_scores_df['Condition']\n            .isin(['baseline', 'only-labeled'])]\n        ), \n        aes(x = 'Model', y = 'support_pct', fill = 'Condition'))\n        + geom_col(position = 'identity')\n        + scale_x_discrete(limits=model_order_by_abstention)\n        + scale_fill_manual(\n            values={'baseline': '#878f99', 'only-labeled': '#1f77b4'},\n            name=\"Support\",\n            labels={'baseline': 'all records', 'only-labeled': 'labeled records'}\n        )\n        + labs(x=\"LLM Model\", y=\"Proportion\", title=\"Support (Positive Labels) Before and After Absention Masking\", fill = \"Support\")\n        + facet_wrap('Finding')\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(10, 10))\n\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:34.093711Z","iopub.execute_input":"2026-09-17T20:14:34.094028Z","iopub.status.idle":"2026-09-17T20:14:35.948209Z","shell.execute_reply.started":"2026-09-17T20:14:34.093998Z","shell.execute_reply":"2026-09-17T20:14:35.947167Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Just a basic gut check, with abstentions we can lose signal if/when we mask - overall the impact is modest on the labeled set except for Synovitis, PF_OA, and in some cases   \n  Effusion.","metadata":{}},{"cell_type":"code","source":"criteria_scores_df['Condition'] = pd.Categorical(criteria_scores_df['Condition'], categories = ['baseline', 'all-filled', 'only-labeled'], ordered = True)\n(\n    ggplot(\n        (\n            criteria_scores_df[['Condition', 'Model', 'f1-score', 'precision', 'recall', 'balanced_accuracy_score', 'npv', 'specificity']]\n                .groupby(['Condition', 'Model'])\n                .mean()\n                .reset_index()\n                .melt(id_vars = ['Condition', 'Model'], var_name = 'Metric', value_name = 'Value')\n        ),\n        aes(x = 'Model', y = 'Value', fill = 'Condition')\n    )\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Metric', ncol = 3)\n        + scale_x_discrete(limits = model_order_by_abstention)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'all-filled': '#daa3e9', 'only-labeled': '#1f77b4', 'baseline': '#878f99'}, \n            name=\"Condition\", \n            labels = ['Baseline', 'Updated - Abstains Filled (Neg)', 'Updated - Abstains Masked']\n        )\n        + labs(x=\"LLM Model\", y=\"Macro Average\", title=\"Overall Performance by Model and Metric\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(10, 5))\n\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:35.949474Z","iopub.execute_input":"2026-09-17T20:14:35.949801Z","iopub.status.idle":"2026-09-17T20:14:36.884124Z","shell.execute_reply.started":"2026-09-17T20:14:35.949761Z","shell.execute_reply":"2026-09-17T20:14:36.883089Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Overall a pretty modest effect of a pretty big change in conditions (addition of significant additional definitional detail for class labeling and the ability to abstain)\n- Modest gains in precision at the macro average level are encouraging likely not statistically significant - at least moving in the right direction if you believe the move\n- Similarity between the two updated conditions is not unexpected: when abstains are filled with negative values they are filled with the most prevalent label\n- Some trade-off in recall, but that is less important given the goal and the fact that precision moved in the right direction\n- **Still no real obvious difference overall across a wide range of models, at least in terms of macro avg across multiple metrics**\n - Note that the fact that OpenAI models are abstaining on up to 27% of cells is a difference, particularly given that doesn't appear translate here to broad gains\n   on precision or NPV","metadata":{}},{"cell_type":"code","source":"(\n    ggplot(criteria_scores_df, aes(x = 'Model', y = 'precision', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = model_order_by_abstention)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'all-filled': '#daa3e9', 'only-labeled': '#1f77b4', 'baseline': '#878f99'}, \n            name=\"Condition\", \n            labels = ['Baseline', 'Updated - Abstains Filled (Neg)', 'Updated - Abstains Masked']\n        )\n        + labs(x=\"LLM Model\", y=\"Precision\", title=\"Positive Predictive Value (Precision) by Condition, Model, and Metric\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:36.885249Z","iopub.execute_input":"2026-09-17T20:14:36.885553Z","iopub.status.idle":"2026-09-17T20:14:38.467674Z","shell.execute_reply.started":"2026-09-17T20:14:36.885524Z","shell.execute_reply":"2026-09-17T20:14:38.466776Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Consistent numerical gains in PPV for ACL, Effusion, MCL, PF_OA, and perhaps Medial_OA, without clear or consistent trends by model\n- Synovitis is strikingly flat; this is one of the markers for which we may have already hit a ceiling governed more by\n  agreement between the original and indepenendent MRI readers.  Others that are similar but less striking include Fracture, Lateral_OA, and Contusion.\n","metadata":{}},{"cell_type":"code","source":"(\n    ggplot(criteria_scores_df, aes(x = 'Model', y = 'npv', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = model_order_by_abstention)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'all-filled': '#daa3e9', 'only-labeled': '#1f77b4', 'baseline': '#878f99'}, \n            name=\"Condition\", \n            labels = ['Baseline', 'Updated - Abstains Filled (Neg)', 'Updated - Abstains Masked']\n        )\n        + labs(x=\"LLM Model\", y=\"NPV\", title=\"Negative Predictive Value (NPV) by Condition, Model, and Metric\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:38.468756Z","iopub.execute_input":"2026-09-17T20:14:38.469028Z","iopub.status.idle":"2026-09-17T20:14:40.015386Z","shell.execute_reply.started":"2026-09-17T20:14:38.469001Z","shell.execute_reply":"2026-09-17T20:14:40.014475Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- So here's the first difference that is at least potentially significant - the new conditions appear to cause a regression on NPV for Effusion,\n  which is not a small deal given that it is the most prevalent finding in the labeled data.  Given the prevalence we are probably calling at least some True Positives\n  Negative based on not meeting the stricter criterion.\n- In general, and not just related to NPV, it seems unlikely that abstentions are helping that much - Anthropic models generally abstain on few cells while OpenAI models abstain on\n  on many more, without any clear advantage on either side.  One lever may be helping the models reason more consistently on what to abstain from.  Another lever that we haven't probed is whether filtering on certainty would help.\n\n## Next Steps\n- On a few classes, (Synovitis, Contusion, Lateral OA, and Medial OA) there is probably some prompt tuning to be done or attempted\n- Beyond that there may be additional value that can be squeezed here, but its not going to come from easy stuff\n  - *Option 1*: Across the LLMs, metrics look similar, but agreement is varies meaningfully - is there an ensemble that could perform better than any single model alone?  Realistically, for 4000+ API calls and the limited variation on metrics by model, I'm probably not looking to do more than 1 model from OpenAI and one from Anthropic;\n  prefereably cheaper ones.  Happily Luna, Terra, and Haiku, are the ones that have the least agreement and also cost the least, meaning there is at least some potential to capture\n  productive differences in labeling tendencies at reasonable cost\n  - *Option 2*: We might be able to squeeze more out of prompt engineering as above - either with dspy prompt hacks, or having astra/fable evalute gaps in reasoning and tweaking manually\n\n## Decision\n- So far we've seen very little effect of model complexity and cost varies greatly.  Going forward we'll limit further testing to luna, haiku, and terra, with a preference to take forward an OpenAI model if possible, due to latency and dspy flakiness, unless there is a clear reason to include haiku.","metadata":{}},{"cell_type":"markdown","source":"## 7. Potential for Uplift from a Lower Cost Ensemble \n- Among lower cost models (terra, haiku, luna) - the best potential to optimize cost and information \n  probably sits in haiku and luna\n- Create 'oracle' labels for a hypothetical model that could always choose the best label between the two \n- Compare to baseline performance, as well as performance in round 2 (adding abstention and criteria)","metadata":{}},{"cell_type":"code","source":"def apply_oracle_lables(labels_df: pd.DataFrame, pred_cols: list = MODEL_KEYS, ref_col: str = 'gold') -> int:\n    def _make_oracle_label(row):\n        if row.gold_agreement:\n            return row.gold\n        else:\n            # Prefer abstention over errors\n            if row.abstention:\n                return np.nan\n            elif row.gold == 1:\n                return 0\n            else:\n                return 1\n    df = labels_df.copy()\n    df['gold_agreement'] = df[pred_cols].eq(df[ref_col], axis=0).any(axis=1)\n    df['abstention'] = df[['luna', 'haiku']].isna().any(axis = 1)\n    df['oracle'] = df.apply(_make_oracle_label, axis = 1)\n    return df['oracle'].astype('Int64')\n\noracle_labels_df = pd.merge(\n    (criteria_preds_long_df\n        .drop(columns = ['rationale', 'Undetermined', 'confidence'])\n        .pivot(index = ['StudyInstanceUID', 'Finding'], columns = 'Model', values = 'annotation')\n        .reset_index()),\n    (gold_df.drop(columns = ['Report', 'Is_English', 'Num_Findings', 'Above_Median_Findings', 'Language'])\n        .melt(id_vars = 'StudyInstanceUID', var_name = 'Finding', value_name = 'gold')),\n    on = ['StudyInstanceUID', 'Finding'],\n    how = 'left'\n)\noracle_labels_df['oracle'] = apply_oracle_lables(oracle_labels_df, pred_cols = ['haiku', 'luna'])\noracle_labels_df = (oracle_labels_df[['StudyInstanceUID', 'Finding', 'oracle']]\n    .pivot(index = 'StudyInstanceUID', columns = 'Finding', values = 'oracle')\n    .reset_index())\n\noracle_scores = pd.DataFrame(get_scores_by_class(gold_df, oracle_labels_df)).rename_axis('metric').reset_index(level = 'metric')\noracle_scores['Model'] = 'Oracle Combo'  \noracle_scores = (oracle_scores.melt(id_vars = ['Model', 'metric'], var_name = 'Finding', value_name = 'Value')\n    .pivot(index = ['Model', 'Finding'], columns = 'metric', values = 'Value')\n    .reset_index()\n)\noracle_scores['support_pct'] = oracle_scores['support'] / baseline_scores_df['support'] \noracle_scores['Condition'] = 'oracle'\noracle_scores = pd.concat([oracle_scores, baseline_scores_df, criteria_masked_scores_df])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:40.016596Z","iopub.execute_input":"2026-09-17T20:14:40.016934Z","iopub.status.idle":"2026-09-17T20:14:40.373412Z","shell.execute_reply.started":"2026-09-17T20:14:40.016879Z","shell.execute_reply":"2026-09-17T20:14:40.372431Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"oracle_scores['Condition'] = pd.Categorical(oracle_scores['Condition'], categories = ['only-labeled','baseline','oracle'], ordered = True)\n(\n    ggplot(\n        (\n            oracle_scores[['Condition', 'Model', 'f1-score', 'precision', 'recall', 'balanced_accuracy_score', 'npv', 'specificity']]\n                .groupby(['Condition', 'Model'])\n                .mean()\n                .reset_index()\n                .melt(id_vars = ['Condition', 'Model'], var_name = 'Metric', value_name = 'Value')\n        ),\n        aes(x = 'Model', y = 'Value', fill = 'Condition')\n    )\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Metric', ncol = 3)\n        + scale_x_discrete(limits = model_order_by_abstention.extend('Oracle Combo'))\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'oracle': '#1f77b4', 'only-labeled':'#daa3e9', 'baseline': '#878f99'}, \n            name=\"Condition\", \n            labels = ['Abstains Masked', 'Baseline', 'Oracle']\n        )\n        + labs(x=\"LLM Model\", y=\"Macro Average\", title=\"Overall Performance by Model and Metric\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(10, 5))\n\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:40.374466Z","iopub.execute_input":"2026-09-17T20:14:40.374727Z","iopub.status.idle":"2026-09-17T20:14:41.837819Z","shell.execute_reply.started":"2026-09-17T20:14:40.374699Z","shell.execute_reply":"2026-09-17T20:14:41.836914Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(oracle_scores, aes(x = 'Model', y = 'precision', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = model_order_by_abstention.extend('Oracle Combo'))\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'oracle': '#1f77b4', 'only-labeled':'#daa3e9', 'baseline': '#878f99'}, \n            name=\"Condition\", \n            labels = ['Abstains Masked', 'Baseline', 'Oracle']\n        )\n        + labs(x=\"LLM Model\", y=\"Precision\", title=\"Positive Predictive Value (Precision) by Condition, Model, and Metric\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:41.839109Z","iopub.execute_input":"2026-09-17T20:14:41.839571Z","iopub.status.idle":"2026-09-17T20:14:43.459992Z","shell.execute_reply.started":"2026-09-17T20:14:41.839514Z","shell.execute_reply":"2026-09-17T20:14:43.458978Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(oracle_scores, aes(x = 'Model', y = 'npv', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = model_order_by_abstention.extend('Oracle Combo'))\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values={'oracle': '#1f77b4', 'only-labeled':'#daa3e9', 'baseline': '#878f99'}, \n            name=\"Condition\", \n            labels = ['Abstains Masked', 'Baseline', 'Oracle']\n        )\n        + labs(x=\"LLM Model\", y=\"NPV\", title=\"Negative Predictive Value (NPV) by Condition, Model, and Metric\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:43.461126Z","iopub.execute_input":"2026-09-17T20:14:43.461534Z","iopub.status.idle":"2026-09-17T20:14:45.081922Z","shell.execute_reply.started":"2026-09-17T20:14:43.461502Z","shell.execute_reply":"2026-09-17T20:14:45.080818Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Hmm.  There might be some minimal potential here, but not a lot, which is not terribly unexpected. \n- At least with the current prompt, this probably is not enough to justify the expense of labeling 4000+ records with two LLMs.\n\n## Decision\n- Further testing with 'haiku', 'luna', and 'terra' only","metadata":{}},{"cell_type":"markdown","source":"## 8. LLM-Assisted Prompt Optimization\n- We could try the dspy tricks, but I'm pretty unoptimistic this will make a difference given the negligible effect of adding the criteria\n- Let's try a single shot handroll informed by Opus 5 looking across all of Terra and Lunas predictions","metadata":{}},{"cell_type":"code","source":"def map_confusion(row) -> str:\n    if pd.isna(row.annotation):\n        return 'ND' \n    if row.annotation == row.gold_label:\n        return 'TP' if row.gold_label == 1 else 'TN'\n    else:\n        return 'FN' if row.gold_label == 1 else 'FP'\n\nllm_data_evaluation_df = pd.merge(\n    criteria_preds_long_df[criteria_preds_long_df['Model'].isin(['luna', 'terra'])],\n    (gold_df[gold_df['StudyInstanceUID'].isin(gold_train_ids['StudyInstanceUID'])]\n        .drop(columns = ['Language', 'Is_English', 'Num_Findings', 'Above_Median_Findings'])\n        .melt(id_vars = ['StudyInstanceUID', 'Report'], var_name = 'Finding', value_name = 'gold_label')),\n    on = ['StudyInstanceUID', 'Finding'],\n    how = 'left'\n)\nllm_data_evaluation_df['confusion_label'] = llm_data_evaluation_df.apply(map_confusion, axis = 1)\nllm_data_evaluation_df = (llm_data_evaluation_df[['StudyInstanceUID', 'Model', 'Report', 'Finding', 'annotation', 'confidence', 'rationale', 'confusion_label']]\n   .pivot(index = ['StudyInstanceUID', 'Report', 'Model'], columns = 'Finding', values = ['annotation', 'confidence', 'rationale', 'confusion_label']))\n\nllm_data_evaluation_df.to_csv('llm_data_evaluation.csv')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:45.08311Z","iopub.execute_input":"2026-09-17T20:14:45.083647Z","iopub.status.idle":"2026-09-17T20:14:45.148125Z","shell.execute_reply.started":"2026-09-17T20:14:45.083615Z","shell.execute_reply":"2026-09-17T20:14:45.147209Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Knee MRI Report Annotation — Prompt Refinement Brief for Claude Code\n\n## Task\n\nAnalyze the patterns of LLM annotations, rationales, and confidence levels across the 12 classes, based on information derived from the `report` column.\n\nNote that the gold labels are derived from an independent read of the MRIs, **not** the reports, so 100% agreement is not possible.\n\n## Goals\n\nHelp refine a basic prompt instruction to:\n\n1. **Guide LLM reasoning to optimize TP and TN rates generally**, by providing general principles on how to label effectively and when to abstain.\n2. **Pay special attention to opportunities to improve LLM reasoning** on Baker (Baker's cyst), Effusion, Synovitis, Contusion, Lateral_OA, and Fracture.\n3. **Do not hardcode information that will help classify specific reports.** The effectiveness of the new prompt will be tested on data you don't see, so this won't be helpful.\n\n## Current Prompt Instructions\n\n```python\nclass KneeMRIOutputCriteria(dspy.Signature):\n    \"\"\"\n    - Act as an expert musculoskeletal radiologist to annotate a knee MRI radiology report for each class of finding:\n      - Use the definitions for each class of finding when making annotations\n      - **Classes of findings**: anterior cruciate ligament (ACL) tear, medial collateral ligament (MCL)\n        tear, medial meniscus tear, lateral meniscus tear, medial tibiofemoral compartment osteoarthritis (Medial_OA),\n        lateral tibiofemoral compartment osteoarthritis (Lateral_OA), patellofemoral compartment osteoarthritis (PF_OA),\n        effusion, synovitis, Bakers' (popliteal) cyst, contusion, and/or fracture\n      - *Annotations for each class*:\n        - **PRESENT**: Based on the text of the radiology report and the class definition, the finding is clearly identified or it can be\n          reasonably inferred if not named directly.\n        - **ABSENT**: Based on the text of the radiology report and the class definition, the finding is clearly excluded or its exclusion can be\n          reasonably inferred if not named directly.\n        - **UNDETERMINED**: Based on the text of the radiology report and the class definition, a determination of **PRESENT** or **ABSENT**\n          cannot be made.\n      - Each class takes an object with an annotation, a confidence rating, and a short rationale citing the report language used.\n    \"\"\"\n    report: str = dspy.InputField()\n    classification: KneeFindingsCriteria = dspy.OutputField()\n```\n\n## Constraints on the Schema\n\nThe LLMs are provided the criteria used by the independent reader to label each class as part of a Pydantic BaseModel for structured outputs. **Do not duplicate this information in any prompt**, although you may suggest slight modifications to the schema (e.g., make a definition less specific).\n\nOverall the schema works well. Aside from specific changes requested below, structural suggestions are out of scope.\n\n## Pydantic BaseModel and Components\n\n```python\nclass ClassificationLabel(Enum):\n    PRESENT = 1\n    ABSENT = 0\n    UNDETERMINED = None\n\n\nclass ConfidenceRating(Enum):\n    VERY_HIGH = 5\n    HIGH = 4\n    MODERATE = 3\n    LOW = 2\n    VERY_LOW = 1\n    NO_CONFIDENCE = 0\n\n\nclass FindingAnnotation(TypedDict):\n    annotation: Annotated[\n        ClassificationLabel,\n        Field(\n            description=\n\"\"\"PRESENT, ABSENT, or UNDETERMINED for this finding, per the finding definition.\"\"\"\n        ),\n    ]\n    confidence: Annotated[\n        ConfidenceRating,\n        Field(\n            description=\n\"\"\"Confidence in the assigned annotation, based on the explicitness and internal consistency\nof the report text.\"\"\"\n        ),\n    ]\n    rationale: Annotated[\n        str,\n        Field(\n            description=\n\"\"\"Short justification (in English, regardless of the report language) for the assigned\nannotation and confidence.\"\"\"\n        ),\n    ]\n\n\nclass ReportFindingsCriteria(TypedDict):\n    ACL: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A high-grade partial or full-thickness tear of the anterior cruciate ligament, meaning complete discontinuity\nof the ligament, or more than 50 percent of fibers disrupted, with or without secondary signs such as characteristic\npivot-shift bone contusions. Mild signal change, degeneration, or thickening without discontinuity is graded negative.\"\"\"\n        ),\n    ]\n    MCL: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A high-grade partial or complete acute tear of the medial collateral ligament, with disrupted fibers\nand edema within and adjacent to the ligament. Low-grade sprains and chronic or remote stress changes are graded negative.\"\"\"\n        ),\n    ]\n    Medial_Meniscus: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the medial meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Lateral_Meniscus: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the lateral meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Medial_OA: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the medial compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Lateral_OA: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the lateral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    PF_OA: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the patellofemoral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Effusion: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large amount of fluid distending the joint.\"\"\"\n        ),\n    ]\n    Synovitis: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"Inflammation and thickening of the synovial lining of the joint.\"\"\"\n        ),\n    ]\n    Baker: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A moderate or large fluid collection in the characteristic Baker (popliteal) cyst location behind the knee.\"\"\"\n        ),\n    ]\n    Contusion: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"A bone contusion, seen as bone marrow edema-like signal from impact, without a discrete fracture line.\"\"\"\n        ),\n    ]\n    Fracture: Annotated[\n        FindingAnnotation,\n        Field(\n            description=\n\"\"\"An acute cortical break or fracture line.\"\"\"\n        ),\n    ]\n```\n\n## Deliverables\n\n- Recommended modifications to the prompt instructions.\n- Slight wording changes to the schema, where warranted.\n- A new `Enum` to capture rationales more consistently (e.g., categories such as `NO_MENTION = \"report does not mention classification at all\"`).\n\nReturn all of the above in a well-formatted markdown document. Agents may be used to perform the analysis if helpful.\n","metadata":{}},{"cell_type":"markdown","source":"# Knee MRI Report Annotation — Prompt Refinement Recommendations (Claude Opus 5)\n\nBased on 38 reports × 2 models (`luna`, `terra`) = 912 class-level annotations in\n`llm_data_evaluation.csv`.\n\n## 1. What the data shows\n\n### 1.1 Baseline performance\n\n| Class | TP | TN | FP | FN | ND | Acc (incl. ND as miss) |\n|---|---|---|---|---|---|---|\n| ACL | 32 | 39 | 0 | 2 | 3 | 0.93 |\n| MCL | 16 | 53 | 1 | 1 | 5 | 0.91 |\n| Medial_Meniscus | 30 | 36 | 4 | 4 | 2 | 0.87 |\n| Medial_OA | 12 | 44 | 1 | 6 | 13 | 0.74 |\n| PF_OA | 11 | 45 | 0 | 8 | 12 | 0.74 |\n| Contusion | 22 | 32 | 13 | 6 | 3 | 0.71 |\n| Effusion | 23 | 31 | 3 | 13 | 6 | 0.71 |\n| Fracture | 8 | 45 | 11 | 8 | 4 | 0.70 |\n| Lateral_Meniscus | 20 | 33 | 3 | 9 | 11 | 0.70 |\n| Lateral_OA | 6 | 47 | 5 | 5 | 13 | 0.70 |\n| Baker | 8 | 40 | 4 | 8 | 16 | 0.63 |\n| Synovitis | 14 | 23 | 6 | 13 | 20 | 0.49 |\n\nOverall: TP 202, TN 468, FP 51, FN 83, UNDETERMINED 108.\n\n### 1.2 Abstention is the single largest loss of accuracy\n\n108 cells (11.8%) are UNDETERMINED, and an UNDETERMINED can never be a TP or TN.\nFor 50 of those 108 cells, the partner model produced a scorable label on the same\nreport, which recovers the gold value. Of those 50, **35 (70%) were gold-negative**.\nConverting abstentions to ABSENT would therefore have produced roughly 2.3 TN for\nevery 1 FN it created.\n\nThe split depends on *why* the model abstained:\n\n| Abstention reason (from rationale text) | n recovered | gold POS | gold NEG |\n|---|---|---|---|\n| Finding not mentioned / region not assessed | 31 | 7 (23%) | 24 (77%) |\n| Finding mentioned but severity/size unquantified | 19 | 8 (42%) | 11 (58%) |\n\nSilence is informative and should resolve to ABSENT. An unquantified positive\nmention is close to a coin flip and should resolve toward PRESENT for the classes\nwhere the reader's threshold is lenient (Effusion: 4 of 6 unquantified-mention\nabstentions were gold-positive).\n\n### 1.3 Confidence is badly calibrated\n\n| Confidence | correct (TP+TN) | incorrect (FP+FN) | ND |\n|---|---|---|---|\n| VERY_HIGH (5) | 489 | 73 | 0 |\n| HIGH (4) | 134 | 37 | 8 |\n| MODERATE (3) | 47 | 24 | 44 |\n| LOW (2) | 0 | 0 | 56 |\n\nVERY_HIGH carries 73 errors, including 40 of the 51 FPs and 33 of the 83 FNs.\nThe models rate confidence on **how explicit the report text is**, which is\nexactly the current schema instruction. That is the wrong target: the report text\ncan be perfectly explicit while still disagreeing with an independent image read.\nLOW is used exclusively as a synonym for abstention, so the bottom of the scale\ncarries no information about assertions.\n\n### 1.4 Threshold literalism is the dominant FN mechanism\n\nSplitting ABSENT calls by whether the rationale invokes a size/grade threshold:\n\n| ABSENT call type | TN | FN | FN rate |\n|---|---|---|---|\n| Rationale cites a size/severity threshold (\"small\", \"mild\", \"below the threshold\") | 58 | 27 | **32%** |\n| Rationale cites absence or a normal statement | 410 | 56 | 12% |\n\nAll 13 Effusion FNs and 6 of 8 Baker FNs are of the first type. The reader's\nimage-based threshold is systematically more permissive than the reporting\nradiologist's adjective.\n\n### 1.5 Class-specific mechanisms\n\n**Baker.** Reported size does not predict the gold label. Gold-positive examples\ninclude a 17 × 12 mm cyst called \"small\" and a 5 mm cyst; gold-negative examples\ninclude a 22 × 12 × 27 mm cyst and a 2 × 1.5 × 3.5 cm cyst. Across all cells where\na popliteal cyst was named at all, 14 of 18 were gold-positive (78%). Meanwhile,\n16 cells abstained on silence, and every recoverable one was gold-negative.\nThe size qualifier is noise; the *naming* is the signal.\n\n**Effusion.** Unqualified \"effusion\" / \"derrame\" / \"Gelenkerguss\" with no amount\n→ gold-positive in 4 of 6 recovered cases. \"Mild\"/\"small\"/\"minimal\" → 22 TN vs\n13 FN, i.e. genuinely mixed and not resolvable from text. Explicit \"no effusion\"\n→ reliably TN.\n\n**Synovitis.** Any named synovitis is a good positive signal (14 TP vs 6 FP, 70%\nprecision) and the severity qualifier is irrelevant: \"minimal synovitis\" appears\namong both TPs and FPs. Silence is weakly negative — 36% of silent reports were\ngold-positive — so ABSENT is the right call but a low-confidence one. The current\nbehavior of abstaining on silence (20 cells, the highest of any class) is the\nsingle largest contributor to this class's 0.49 accuracy.\n\n**Contusion and Fracture are being treated as mutually exclusive; the gold labels\nare not.** In gold, 6 of 9 fracture-positive reports are also contusion-positive.\nAll 6 Contusion FNs come from reasoning of the form \"the marrow edema is\nattributed to an impaction/insufficiency fracture, therefore it is not a\ncontusion.\"\n\n**Fracture over-calls sub-cortical entities.** All 11 FPs are reports naming\n\"microtrabecular insufficiency fracture,\" \"subcortical fracture,\" \"impaction\ninjury,\" or \"osteochondral fracture.\" The gold criterion is a discrete cortical\nbreak. Gold-positive fractures were tibial eminence/spine fractures, Segond\nfracture, avulsion fracture, depressed cortical contour, and an explicitly\ndescribed hairline fracture line.\n\n**Lateral_OA.** 13 abstentions, nearly all because the report gives no lateral\ncompartment cartilage assessment; 4 of 5 recoverable were gold-negative. FPs come\nfrom accepting \"grade 2–3 chondropathy,\" \"nearly full-thickness,\" or a bare\n\"OA femorotibial lateral\" as qualifying. TPs are uniformly grade 3–4, \"complete\ncartilage loss,\" or \"advanced.\" For the OA classes, unlike Baker and Synovitis,\nthe severity qualifier does carry signal and should be respected.\n\n---\n\n## 2. Recommended prompt instructions\n\nReplace the `KneeMRIOutputCriteria` docstring with the following. It adds a\ndecision procedure, an abstention policy, an anti-correlation rule, and a\nconfidence definition, without encoding any report-specific content.\n\n```python\nclass KneeMRIOutputCriteria(dspy.Signature):\n    \"\"\"\n    Act as an expert musculoskeletal radiologist annotating a knee MRI radiology report for\n    each class of finding. Use the definition supplied with each class when making annotations.\n\n    WHAT YOU ARE PREDICTING\n    The reference labels come from an independent review of the MRI images, not from this\n    report. Your task is therefore not to summarize the report, but to predict what an\n    independent image reader would have concluded, using the report as evidence. Reports are\n    written for clinical communication: they under-report findings incidental to the referral\n    question, omit findings the referrer did not ask about, and use severity adjectives\n    inconsistently across institutions and languages. Weigh the report accordingly.\n\n    DECISION PROCEDURE\n    For each class, work through these in order:\n      1. Direct statement. Does the report name the finding, affirm it, or deny it?\n      2. Equivalent description. Does it describe the finding's substance under different\n         wording, in another language, or as part of a composite diagnosis? Named entities that\n         subsume the finding count as affirmations.\n      3. Regional negation. Does a normal or negative statement about the relevant structure or\n         compartment cover this finding? Global normality statements and structure-specific\n         negations are evidence of absence, not merely of silence.\n      4. Inference from related findings. Does another reported finding make this one likely or\n         unlikely, given normal co-occurrence of knee pathology?\n      5. Silence. If none of the above applies, apply the abstention policy below.\n    Give the impression/conclusion precedence over the body when they conflict, and say so in\n    the rationale.\n\n    ANNOTATION POLICY\n    - PRESENT: the report affirms the finding, describes its substance, or supports it strongly\n      enough that an independent image reader would more likely than not have recorded it.\n    - ABSENT: the report denies the finding, describes the relevant structures as normal, or is\n      silent about a finding that the report's scope would ordinarily have covered.\n    - UNDETERMINED: reserve this for the rare case where the report supplies genuinely\n      contradictory evidence that cannot be resolved, or where the relevant region was not\n      imaged or was explicitly non-diagnostic. Silence alone is not grounds for UNDETERMINED.\n      An UNDETERMINED can never be scored correct, so an annotation you hold at 50/50 is still\n      worth committing; commit to the more likely label and record low confidence instead.\n\n    HOW TO HANDLE SILENCE\n    A report that does not mention a finding is usually a report in which that finding was\n    absent, so silence defaults to ABSENT. Weight this default by how routinely the finding is\n    reported when present: findings that radiologists name whenever they see them (discrete\n    ligament tears, fractures, cysts, effusions) make silence strong evidence of absence;\n    findings that are frequently observed but omitted as incidental (synovial change,\n    compartment-specific cartilage grading in a trauma report) make silence weak evidence of\n    absence. Annotate ABSENT in both cases, and let confidence carry the difference.\n\n    HOW TO HANDLE SEVERITY QUALIFIERS\n    Severity adjectives (\"mild\", \"small\", \"minimal\", \"slight\", \"leve\", \"gering\") are the single\n    largest source of false negatives in this task. They are the reporting radiologist's\n    subjective grading on a different modality view, and they frequently sit on the opposite\n    side of the reference reader's threshold. Apply these rules:\n      - Where a class definition has a size or severity threshold, that threshold is the\n        reference reader's, not the reporting radiologist's. Do not resolve to ABSENT on the\n        strength of a diminutive adjective alone when the finding itself is affirmed. Weigh the\n        objective content of the description (measurements, extent, number of regions, secondary\n        signs) above the adjective.\n      - Distinguish a diminutive qualifier from a negation. \"Mild effusion\" affirms the finding;\n        \"no significant effusion\" and \"no effusion\" negate it.\n      - A finding that is named but not quantified should be graded PRESENT, not UNDETERMINED.\n        Radiologists tend to name findings that are conspicuous.\n      - The exception is the cartilage/osteoarthritis classes, where the grading vocabulary is\n        standardized (Outerbridge/ICRS grade, \"full-thickness\", \"advanced\", \"complete cartilage\n        loss\") and does track the reference threshold. There, respect the reported grade, and\n        treat low-grade chondral change, fissuring, thinning, an unqualified diagnosis of\n        \"osteoarthritis\", or isolated osteophytes as not meeting the definition.\n\n    FINDINGS ARE NOT MUTUALLY EXCLUSIVE\n    Annotate each class independently against its own definition. In particular, a report that\n    attributes an abnormality to one entity does not thereby exclude an overlapping class: bone\n    marrow edema accompanying a fracture still satisfies a marrow-edema definition, and\n    traumatic and degenerative explanations for the same signal can both be recorded. Do not\n    suppress one class because another class explains the same sentence. Conversely, an entity\n    named in the report does not automatically satisfy a class whose definition is narrower than\n    the named entity; check the definition's specific requirement (for example, a required\n    cortical break) against what was actually described, and grade the finding under whichever\n    class the description actually matches.\n\n    CONFIDENCE\n    Rate confidence in the probability that an independent image reader assigned the same label,\n    not in how explicit the report text is. Lower confidence when the report is explicit but the\n    class definition turns on a threshold the report does not let you evaluate, when you are\n    inferring from silence, when the finding is one radiologists commonly omit, or when the\n    report's referral question would not have prompted an assessment of this class. Reserve\n    VERY_HIGH for a direct affirmation or a direct negation of a finding whose definition the\n    report clearly satisfies or clearly fails. Use LOW and VERY_LOW for committed annotations\n    you hold weakly; they are not reserved for UNDETERMINED.\n\n    RATIONALE\n    Each class takes an object with an annotation, a confidence rating, an evidence category,\n    and a short rationale. The rationale must quote or closely paraphrase the specific report\n    language relied on, or state that the report is silent. It must be consistent with the\n    evidence category and with the annotation.\n    \"\"\"\n    report: str = dspy.InputField()\n    classification: KneeFindingsCriteria = dspy.OutputField()\n```\n\n### 2.1 Why each block is there\n\n| Block | Failure it targets |\n|---|---|\n| \"What you are predicting\" | The models reason as report summarizers, which caps achievable agreement and drives the threshold literalism in §1.4. |\n| Decision procedure | Regional negation and inference from related findings were used inconsistently; making them explicit steps recovers TNs currently sitting in ND. |\n| Abstention policy | 108 ND cells, ~70% of them gold-negative (§1.2). |\n| Silence handling | Baker (16 ND), Synovitis (20 ND), Lateral_OA (13 ND), Medial_OA (13 ND), PF_OA (12 ND). |\n| Severity qualifiers | 27 FNs from threshold-citing ABSENT calls; Effusion and Baker specifically. |\n| Non-exclusivity | 6 Contusion FNs from fracture-attributed edema; 11 Fracture FPs from sub-cortical entities graded as fractures. |\n| Confidence | 73 errors at VERY_HIGH; LOW/VERY_LOW unused for assertions. |\n\n---\n\n## 3. Recommended schema wording changes\n\nMinimal edits, all narrowing or loosening a definition that the error analysis\nshows is being read at the wrong strictness. The reference reader's criteria are\nunchanged in substance.\n\n**`Effusion`** — the bare threshold invites literal adjective matching.\n> A moderate or large amount of fluid distending the joint. Judge the amount from the described\n> extent and distribution (recess or bursal distention, multi-compartment fluid) rather than from\n> the reporting radiologist's severity adjective alone.\n\n**`Baker`** — reported size does not predict the reference label (§1.5).\n> A fluid collection in the characteristic Baker (popliteal) cyst location behind the knee,\n> excluding trace or incidental fluid.\n\n*(Dropping \"moderate or large\" is the highest-value single change in this\ndocument: 6 of 8 Baker FNs are cysts rejected on reported size.)*\n\n**`Synovitis`** — add the vocabulary variants, drop any implied severity floor.\n> Inflammation or thickening of the synovial lining of the joint, including synovial\n> hypertrophy, synovial proliferation, reactive or chronic synovitis, and synovial\n> nodularity, at any reported severity.\n\n**`Fracture`** — the discriminating requirement is the cortical break.\n> An acute cortical break or discrete fracture line. Trabecular, subchondral, or impaction\n> marrow injury without a described cortical break or fracture line does not qualify and should\n> be considered under the contusion definition instead.\n\n**`Contusion`** — make co-occurrence with fracture explicit.\n> Bone marrow edema-like signal from impact, without a discrete fracture line through that area\n> of edema. Marrow edema in a knee that also contains a fracture still qualifies, provided the\n> edema itself is not described solely as a fracture line.\n\n**`Lateral_OA` / `Medial_OA` / `PF_OA`** — append one sentence to each (they are\notherwise well calibrated; the FPs come from accepting near-threshold grades):\n> Grade 1–2 chondral change, fissuring, thinning, partial-thickness loss, isolated osteophytes,\n> and an unqualified diagnosis of \"osteoarthritis\" or \"gonarthrosis\" do not qualify.\n\nNo changes recommended for `ACL`, `MCL`, `Medial_Meniscus`, or `Lateral_Meniscus`\n(0.87–0.93 accuracy, 8 FPs and 7 FNs combined across all four).\n\n**`confidence` field description** — retarget from text explicitness to label agreement:\n> Probability that an independent reader of the MRI images assigned this same label, expressed\n> on the confidence scale. Lower it when the annotation rests on silence, on an inference, or on\n> a threshold the report does not let you evaluate, even where the report language is explicit.\n\n---\n\n## 4. New Enum: `EvidenceBasis`\n\nCaptures *why* an annotation was made, in a form that is countable across a run.\nFive categories, each describing the **kind** of report evidence used. The\ndirection (present/absent) stays in the `annotation` field, so the auditable unit\nis the (basis, annotation) pair.\n\n```python\nclass EvidenceBasis(Enum):\n    \"\"\"The kind of report evidence the annotation rests on. Independent of the\n    annotation: STATED, DESCRIBED, and THRESHOLD can each support PRESENT or ABSENT.\"\"\"\n\n    STATED = \"the report names the finding and either asserts or denies it\"\n    DESCRIBED = \"the finding is not named; the annotation follows from described morphology, or from the relevant structure being described as normal\"\n    THRESHOLD = \"the finding is affirmed, and the annotation turns on whether its reported severity, size, or extent meets the definition's threshold\"\n    NOT_MENTIONED = \"the report evaluates the relevant region but does not mention this finding\"\n    NOT_EVALUABLE = \"the relevant region was not imaged or not assessed, or the report's statements about it are contradictory or too equivocal to resolve\"\n```\n\nAdd to `FindingAnnotation`, ahead of `rationale`:\n\n```python\n    evidence_basis: Annotated[\n        EvidenceBasis,\n        Field(\n            description=\n\"\"\"The kind of report evidence this annotation rests on. Select the most specific applicable\ncategory: prefer STATED or DESCRIBED over THRESHOLD, and use NOT_MENTIONED only when neither\napplies.\"\"\"\n        ),\n    ]\n```\n\nThe `NOT_MENTIONED` / `NOT_EVALUABLE` split is the one that has to be kept.\nCollapsing them into a single \"not addressed\" category merges the case that must\nresolve to ABSENT (§1.2: silence, 77% gold-negative) with the only case that may\nresolve to UNDETERMINED, which is precisely the distinction this prompt revision\nis trying to enforce, and it makes the first constraint below uncheckable.\n\n### 4.1 Consistency constraints worth enforcing downstream\n\nCheckable post hoc without a gold label, so they diagnose prompt drift on the\nheld-out data:\n\n| Constraint | Rationale |\n|---|---|\n| `UNDETERMINED` → basis is `NOT_EVALUABLE` | The only basis the revised prompt admits for abstention (§1.2). |\n| `NOT_MENTIONED` → annotation is `ABSENT` | Silence must resolve. |\n| `NOT_MENTIONED` → confidence ≤ HIGH | Silence-based calls carried 56 FNs. |\n| `THRESHOLD` → confidence ≤ HIGH | 32% error rate on threshold-citing calls (§1.4). |\n| `STATED` + annotation contradicting the report's own wording | Permitted, but should be rare, and the rationale should say why the definition inverts the report's label. |\n\nPer-class `THRESHOLD` and `NOT_MENTIONED` rates are the two run-level metrics to\nwatch: the first should fall for Baker and Effusion, the second should replace\nmost of the current 108 abstentions.\n\n### 4.2 Expected effect\n\nResolving abstentions per §1.2's recovered gold rates (70% of abstentions\ngold-negative) would move overall accuracy from 0.735 to roughly 0.82 on this\nsample, before any gain from the Baker, Contusion, and Fracture changes. That\nfigure extrapolates the 50 recoverable ND cells to all 108 and should be treated\nas an estimate, not a measurement.\n","metadata":{}},{"cell_type":"markdown","source":"## 9. Refined Labeling Program\n- Apply the prompt refinement recommendations (with a some revisions) above to a third version of the labeling program\n- Signature instructions add what the labels actually are (an independent image read, not a report\n  summary), an ordered decision procedure, an abstention policy that resolves silence to ABSENT, a\n  severity-qualifier policy, a non-exclusivity rule, and a confidence definition targeted at label\n  agreement rather than report explicitness\n- Class definitions revised for Effusion, Synovitis, Baker, Contusion, Fracture and the three OA\n  classes; ACL, MCL and the two meniscus classes carry forward unchanged\n- Each annotation now also commits to an `evidence_basis` naming the kind of report evidence it\n  rests on, keeping NOT_MENTIONED (silence, must resolve to ABSENT) separate from NOT_EVALUABLE\n  (the only basis that admits UNDETERMINED)\n- Self-contained: every type is redeclared here rather than reused from the criteria program","metadata":{}},{"cell_type":"code","source":"class ClassificationLabelRefined(Enum):\n    PRESENT = 1\n    ABSENT = 0\n    UNDETERMINED = None\n\n\nclass ConfidenceRatingRefined(Enum):\n    VERY_HIGH = 5\n    HIGH = 4\n    MODERATE = 3\n    LOW = 2\n    VERY_LOW = 1\n    NO_CONFIDENCE = 0\n\n\nclass EvidenceBasis(Enum):\n    \"\"\"The kind of report evidence the annotation rests on. Independent of the\n    annotation: STATED, DESCRIBED, and THRESHOLD can each support PRESENT or ABSENT.\"\"\"\n\n    STATED = \"the report names the finding and either asserts or denies it\"\n    DESCRIBED = \"the finding is not named; the annotation follows from described morphology, or from the relevant structure being described as normal\"\n    THRESHOLD = \"the finding is affirmed, and the annotation turns on whether its reported severity, size, or extent meets the definition's threshold\"\n    NOT_MENTIONED = \"the report evaluates the relevant region but does not mention this finding\"\n    NOT_EVALUABLE = \"the relevant region was not imaged or not assessed, or the report's statements about it are contradictory or too equivocal to resolve\"\n\n\nclass FindingAnnotationRefined(TypedDict):\n    annotation: Annotated[\n        ClassificationLabelRefined,\n        Field(\n            description=\n\"\"\"PRESENT, ABSENT, or UNDETERMINED for this finding, per the finding definition.\"\"\"\n        ),\n    ]\n    confidence: Annotated[\n        ConfidenceRatingRefined,\n        Field(\n            description=\n\"\"\"Confidence in the specific annotation made based on the annotation policy\"\"\"\n        ),\n    ]\n    evidence_basis: Annotated[\n        EvidenceBasis,\n        Field(\n            description=\n\"\"\"The kind of report evidence this annotation rests on. Select the most specific applicable\ncategory: prefer STATED or DESCRIBED over THRESHOLD, and use NOT_MENTIONED only when neither\napplies.\"\"\"\n        ),\n    ]\n    rationale: Annotated[\n        str,\n        Field(\n            description=\n\"\"\"Short justification (in English, regardless of the report language) for the assigned\nannotation and confidence. Quote or closely paraphrase the report language relied on, or state\nthat the report is silent.\"\"\"\n        ),\n    ]\n\n\nclass ReportFindingsRefined(TypedDict):\n    ACL: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"A high-grade partial or full-thickness tear of the anterior cruciate ligament, meaning complete discontinuity\nof the ligament, or more than 50 percent of fibers disrupted, with or without secondary signs such as characteristic\npivot-shift bone contusions. Mild signal change, degeneration, or thickening without discontinuity is graded negative.\"\"\"\n        ),\n    ]\n    MCL: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"A high-grade partial or complete acute tear of the medial collateral ligament, with disrupted fibers\nand edema within and adjacent to the ligament. Low-grade sprains and chronic or remote stress changes are graded negative.\"\"\"\n        ),\n    ]\n    Medial_Meniscus: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the medial meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Lateral_Meniscus: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the lateral meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Medial_OA: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the medial compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Lateral_OA: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the lateral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    PF_OA: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the patellofemoral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Effusion: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"Fluid distending the joint.\"\"\"\n# Removed the size qualification.\n        ),\n    ]\n    Synovitis: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"Inflammation or thickening of the synovial lining of the joint, including synovial hypertrophy, synovial proliferation,\nreactive or chronic synovitis, and synovial nodularity, at any reported severity.\"\"\"\n        ),\n    ]\n    Baker: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"A fluid collection in the characteristic Baker (popliteal) cyst location behind the knee, excluding trace or incidental\nfluid.\"\"\"\n        ),\n    ]\n    Contusion: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"Bone marrow edema-like signal from impact area of edema. Marrow edema in a knee that also contains a fracture still qualifies, \nprovided the edema itself is not described solely as a fracture line.\"\"\"\n# Modified to clarify that a fracture does not per se exclude edema\n        ),\n    ]\n    Fracture: Annotated[\n        FindingAnnotationRefined,\n        Field(\n            description=\n\"\"\"An acute cortical break or discrete fracture line. Trabecular, subchondral, or impaction marrow injury without a described\ncortical break or fracture line does not qualify and should be considered under the contusion definition instead.\"\"\"\n# Modified to clarify when contusion might be more appropriate\n        ),\n    ]\n\n\n# Same 12 names as FINDING_COLUMNS, derived from the schema so the two cannot drift apart\nFINDING_NAMES_REFINED = tuple(ReportFindingsRefined.__annotations__)\nENUM_KEYS_REFINED = {'annotation': ClassificationLabelRefined, 'confidence': ConfidenceRatingRefined}\n\n\ndef _as_int_refined(value, enum_cls: type[Enum]) -> int | None:\n    if value is None:\n        return None\n    if isinstance(value, enum_cls):\n        return None if value.value is None else int(value.value)\n    if isinstance(value, str):\n        name = value.strip().upper()\n        if name in enum_cls.__members__:\n            member = enum_cls[name]\n            return None if member.value is None else int(member.value)\n        value = int(value)\n    member = enum_cls(int(value))\n    return None if member.value is None else int(member.value)\n\n\ndef _as_evidence_basis_name(value) -> str | None:\n    if value is None:\n        return None\n    if isinstance(value, EvidenceBasis):\n        return value.name\n    name = str(value).strip().upper()\n    if name in EvidenceBasis.__members__:\n        return name\n    return EvidenceBasis(str(value)).name\n\n\ndef as_long_records_refined(findings: dict, study_instance_uid: str) -> list[dict]:\n    records = []\n    for name in FINDING_NAMES_REFINED:\n        annotation = findings.get(name) or {}\n        rationale = annotation.get('rationale')\n        records.append({\n            'StudyInstanceUID': str(study_instance_uid),\n            'Finding': name,\n            **{key: _as_int_refined(annotation.get(key), cls) for key, cls in ENUM_KEYS_REFINED.items()},\n            'evidence_basis': _as_evidence_basis_name(annotation.get('evidence_basis')),\n            'rationale': None if rationale is None else str(rationale),\n        })\n    return records\n\n\nclass KneeFindingsRefined(BaseModel):\n    findings: ReportFindingsRefined = Field(\n        description=\n\"\"\"Annotations per class for the knee MRI radiology report\"\"\"\n    )\n\n\nclass KneeMRIOutputRefined(dspy.Signature):\n    \"\"\"\n    Act as an expert musculoskeletal radiologist annotating a knee MRI radiology report for\n    each class of finding. Use the definition supplied with each class when making annotations.=\n\n    DECISION PROCEDURE\n    For each class determine in order whether the report text provides a:\n      1. Does the report provide a clear and direct statement that confirmes the presence \n        (or its absence) according to finding criteria, including any severity thresholds?\n      2. Does the report use equivalent language that describes the finding's presence (or absence) \n         (eg, using clinically related language or as part of a composite diagnosis)?\n      3. Does a normal or negative statement about the relevant structure or\n         compartment that would indicate the finding is absent (eg, the relevant structure is\n         described as normal or without defects)?\n      4. Does another reported finding make this one likely or unlikely, given normal co-occurrence \n         of knee pathology?\n      5. Is the report silent on the finding? if none of the above applies, apply the abstention policy below.\n    Give the impression/conclusion precedence over the body when they conflict, and say so in\n    the rationale.\n\n    ANNOTATION POLICY\n    - PRESENT: the report affirms the finding, describes its substance, or supports it strongly\n      enough that an independent image reader would more likely than not have recorded it.\n    - ABSENT: the report denies the finding, describes the relevant structures as normal, or is\n      silent about a finding that would have normally been described, if present, given the other\n      types of abnormalities present.\n    - UNDETERMINED: reserve this for the rare case where the report supplies genuinely\n      contradictory evidence that cannot be resolved, or where the relevant region was not\n      imaged or was explicitly non-diagnostic. Silence alone is not grounds for UNDETERMINED.\n    \n    HOW TO HANDLE SILENCE\n    - For reports that do not mention a finding, generally suspect that finding is ABSENT\n    - An UNDETERMINED rating may, at times, be most appropriate for findings that are frequently \n      observed but omitted as incidental (synovial change, compartment-specific cartilage grading \n      in a trauma report). In such cases silence is significantly weaker evidence of absence.\n    \n    HOW TO HANDLE SEVERITY QUALIFIERS\n      - Where a class definition has a size or severity threshold, do not resolve to ABSENT on the\n        strength of a diminutive adjective alone when the finding itself is affirmed. Weigh the\n        objective content of the description (measurements, extent, number of regions, secondary\n        signs) more importantly.\n      - Distinguish a diminutive qualifier from a negation. \"Mild effusion\" affirms the finding;\n        \"no significant effusion\" and \"no effusion\" indicate the finding is ABSENT.\n      - A finding that is named but not quantified should be graded generally PRESENT, not UNDETERMINED\n        as Radiologists tend to name findings that are conspicuous.\n      - The exception is the cartilage/osteoarthritis classes, where the grading vocabulary is\n        standardized (Outerbridge/ICRS grade, \"full-thickness\", \"advanced\", \"complete cartilage\n        loss\") and does track the reference threshold. There, respect the reported grade, and\n        treat low-grade chondral change, fissuring, thinning, an unqualified diagnosis of\n        \"osteoarthritis\", or isolated osteophytes as not meeting the definition.\n\n    FINDINGS ARE NOT MUTUALLY EXCLUSIVE\n    Annotate each class independently against its own definition. In particular, a report that\n    attributes an abnormality to one entity does not thereby exclude an overlapping class: bone\n    marrow edema accompanying a fracture still satisfies a marrow-edema definition, and\n    traumatic and degenerative explanations for the same signal can both be recorded. \n\n    CONFIDENCE\n    Confidence is the confidence in the specific annotation according to the policy above; it is\n    NOT confidence that a finding is present.  Annotations of PRESENT, ABSENT, and UNDETERMINED\n    will be made with varying levels of confidence, which is what the determination should reflect.\n\n    RATIONALE\n    The rationale must quote or closely paraphrase the specific report language relied on, or state that the report is silent. \n    It must be consistent with the evidence category and with the annotation.\n    \"\"\"\n    report: str = dspy.InputField()\n    classification: KneeFindingsRefined = dspy.OutputField()\n\n\nclass RefinedClassifier(dspy.Module):\n    def __init__(self):\n        super().__init__()\n        self.classification_module = dspy.ChainOfThought(KneeMRIOutputRefined)\n\n    def forward(self, report: str, study_instance_uid: str) -> dspy.Prediction:\n        result = self.classification_module(report=report)\n\n        return dspy.Prediction(\n            study_instance_uid=study_instance_uid,\n            findings=as_long_records_refined(result.classification.findings, study_instance_uid),\n        )\n\ndef format_examples_refined(df: pd.DataFrame) -> list[dspy.Example]:\n    df_findings = df[FINDING_COLUMNS].astype(int).to_dict(orient = 'records')\n    examples = []\n    for i, row in enumerate(df.itertuples()):\n        example = dspy.Example(\n            study_instance_uid = row.StudyInstanceUID,\n            report = row.Report,\n            findings = df_findings[i]\n        ).with_inputs(\"report\", \"study_instance_uid\")\n        examples.append(example)\n    return examples","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:45.149182Z","iopub.execute_input":"2026-09-17T20:14:45.149491Z","iopub.status.idle":"2026-09-17T20:14:45.183337Z","shell.execute_reply.started":"2026-09-17T20:14:45.149463Z","shell.execute_reply":"2026-09-17T20:14:45.182215Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Runs predictions using the refined prompt across luna and terra, unless the serialized\n# output is available\n\nrefined_preds_load_path = DATA_LOC / 'datasets' / 'joshuaziel' / 'rsna-knee-mri-radiology-report-classifications' / 'raw_prediction_results_v3.json'\nrefined_preds_persist_path = OUTPUTS_LOC / 'raw_prediction_results_v3.json'\n\n# Only luna and terra carry forward from the Section 7 decision\nREFINED_MODEL_KEYS = ['luna', 'terra', 'haiku']\n\nif refined_preds_load_path.is_file():\n    with open(str(refined_preds_load_path), 'r') as f:\n        refined_preds_raw = json.load(f)\n    print(\"Loaded LM predictions from file\")\n\n# Same hand-rolled retry loop as the earlier runs.  Note this will now break for fable unless tool_choice = 'auto' is selected\n# but that will break it for OpenAI\nelse:\n    program = RefinedClassifier()\n    refined_preds_raw = {}\n    for model_key in tqdm(REFINED_MODEL_KEYS, dynamic_ncols = True):\n        model_spec = MODELS[model_key]\n        dspy_lm = dspy.LM(model = model_spec.name, temperature = 1.0, model_type = model_spec.type, api_key = model_spec.key, cache = False)\n        dspy.configure(lm=dspy_lm)\n        model_preds = []\n\n        for row in gold_df.itertuples():\n            success, tries = False, 0\n            while not success and tries < 4:\n                tries += 1\n                try:\n                    prediction = program(study_instance_uid=row.StudyInstanceUID, report=row.Report)\n                    model_preds.extend(prediction.findings)\n                    success = True\n                except Exception as e:\n                    tqdm.write(f'Error predicting labels for {row.StudyInstanceUID}: {e}')\n            if not success:\n                model_preds.extend(as_long_records_refined({}, row.StudyInstanceUID))\n        refined_preds_raw[model_key] = model_preds\n\n    with open(str(refined_preds_persist_path), 'w') as f:\n        json.dump(refined_preds_raw, f, indent = 4)\n\nrefined_preds_long_df = pd.concat(\n    [pd.DataFrame(records).assign(Model=model_key) for model_key, records in refined_preds_raw.items()],\n    ignore_index=True,\n)\nrefined_preds_long_df = refined_preds_long_df[refined_preds_long_df['StudyInstanceUID'].isin(gold_train_ids['StudyInstanceUID'])]\nrefined_preds_long_df['annotation'] = refined_preds_long_df['annotation'].astype(\"Int64\")\nrefined_preds_long_df['Undetermined'] = pd.isna(refined_preds_long_df['annotation'])\nrefined_preds_long_filled_df = refined_preds_long_df.copy()\nrefined_preds_long_filled_df.loc[refined_preds_long_filled_df['Undetermined'], 'annotation'] = -1","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:45.184296Z","iopub.execute_input":"2026-09-17T20:14:45.184584Z","iopub.status.idle":"2026-09-17T20:14:45.234961Z","shell.execute_reply.started":"2026-09-17T20:14:45.184554Z","shell.execute_reply":"2026-09-17T20:14:45.233904Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 10. Re-Evaluation of the Refined Program (haiku, luna, and terra)\n- Do the refined instructions change agreement between the three remaining models?\n- Does abstention fall, as the silence-is-evidence policy intends, and what evidence do the\n  annotations now rest on?\n- How does performance compare to baseline and to the criteria run with abstains masked?\n- Predictions are cached for all labeled records; every evaluation below is restricted to the\n  train ids","metadata":{}},{"cell_type":"code","source":"make_cohens_kappa_matrix_plot(\n    refined_preds_long_filled_df,\n    models = {model_key: MODELS[model_key] for model_key in REFINED_MODEL_KEYS},\n    mid_value = 0.7, # 0.7 might generally be consider to be a moderate level of agreement\n    scale_limits = (0.00, 1.0)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:45.236055Z","iopub.execute_input":"2026-09-17T20:14:45.236389Z","iopub.status.idle":"2026-09-17T20:14:45.57833Z","shell.execute_reply.started":"2026-09-17T20:14:45.236359Z","shell.execute_reply":"2026-09-17T20:14:45.577517Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(refined_preds_long_df, aes(x=\"confidence\"))\n        + geom_histogram(bins = 6, color = \"white\", fill = \"#1f77b4\")\n        + facet_grid(cols = 'Finding', rows = 'Model')\n        + scale_x_continuous(breaks=[0,1,3,5], labels=[\"None\", \"Very Low\", \"Moderate\", \"Very High\"])\n        + theme(figure_size=(15, 4))\n        + labs(x=\"Confidence Level\", y=\"Count, n\", title=\"Frequency Distribution of Confidence Levels by LLM and Class - Refined Program\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:45.579459Z","iopub.execute_input":"2026-09-17T20:14:45.579731Z","iopub.status.idle":"2026-09-17T20:14:48.969359Z","shell.execute_reply.started":"2026-09-17T20:14:45.579704Z","shell.execute_reply":"2026-09-17T20:14:48.968314Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"refined_model_order_by_abstention = refined_preds_long_df[['Model', 'Undetermined']].groupby('Model').sum().sort_values(by = 'Undetermined', ascending = True).index.to_list()\n(\n    ggplot(\n        refined_preds_long_df.groupby(['Model', 'Finding'])['Undetermined'].sum().reset_index(),\n        aes(x = 'Model', y = 'Undetermined')\n    )\n    + facet_wrap('Finding')\n    + geom_col(position=\"dodge\", color=\"white\", fill = \"#1f77b4\")\n    + scale_x_discrete(limits=refined_model_order_by_abstention)\n    + theme(figure_size=(8, 6))\n    + labs(x=\"LLM Model\", y=\"Labeling Abstentions, n\", title=\"Frequency of Labeling Abstention by Class and LLM - Refined Program\")\n    + theme(axis_text_x=element_text(angle=45, hjust=1))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:48.970699Z","iopub.execute_input":"2026-09-17T20:14:48.971015Z","iopub.status.idle":"2026-09-17T20:14:49.899867Z","shell.execute_reply.started":"2026-09-17T20:14:48.970973Z","shell.execute_reply":"2026-09-17T20:14:49.899058Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# The refined program commits to an evidence basis before writing a rationale, so the evidence the\n# annotation rests on is auditable separately from the direction of the annotation\nEVIDENCE_BASIS_ORDER = [evidence_basis.name for evidence_basis in EvidenceBasis]\nEVIDENCE_BASIS_COLORS = {\n    'STATED': '#1f77b4',\n    'DESCRIBED': '#6baed6',\n    'THRESHOLD': '#daa3e9',\n    'NOT_MENTIONED': '#878f99',\n    'NOT_EVALUABLE': '#fdae61',\n}\n(\n    ggplot(\n        refined_preds_long_df.groupby(['Model', 'Finding', 'evidence_basis']).size().reset_index(name = 'n'),\n        aes(x = 'Model', y = 'n', fill = 'evidence_basis')\n    )\n    + facet_wrap('Finding')\n    + geom_col(position = 'fill', color = 'white')\n    + scale_x_discrete(limits = refined_model_order_by_abstention)\n    + scale_fill_manual(values = EVIDENCE_BASIS_COLORS, limits = EVIDENCE_BASIS_ORDER, name = \"Evidence Basis\")\n    + labs(x=\"LLM Model\", y=\"Proportion of Annotations\", title=\"Evidence Basis by Class and LLM - Refined Program\")\n    + theme(axis_text_x=element_text(angle=45, hjust=1),\n        figure_size=(10, 8))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:49.900939Z","iopub.execute_input":"2026-09-17T20:14:49.901257Z","iopub.status.idle":"2026-09-17T20:14:51.223501Z","shell.execute_reply.started":"2026-09-17T20:14:49.9012Z","shell.execute_reply":"2026-09-17T20:14:51.222646Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"refined_preds_wide_df = refined_preds_long_df.pivot(index = ['Model', 'StudyInstanceUID'], columns = 'Finding', values = 'annotation').reset_index()\n\nrefined_baseline_score_parts = []\nfor model_key in REFINED_MODEL_KEYS:\n    model_scores_df = pd.DataFrame(get_scores_by_class(\n        gold_df,\n        baseline_preds_wide_df[baseline_preds_wide_df['Model'] == model_key]\n    )).rename_axis('metric').reset_index(level = 'metric')\n    model_scores_df['Model'] = model_key\n    refined_baseline_score_parts.append(model_scores_df)\nrefined_baseline_scores_df = (\n    pd.concat(refined_baseline_score_parts)\n        .melt(id_vars = ['Model', 'metric'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['Model', 'Finding'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\nrefined_baseline_scores_df['support_pct'] = refined_baseline_scores_df['support'] / refined_baseline_scores_df['support']\nrefined_baseline_scores_df['Condition'] = 'baseline'\n\nrefined_criteria_score_parts = []\nfor model_key in REFINED_MODEL_KEYS:\n    model_scores_df = pd.DataFrame(get_scores_by_class(\n        gold_df,\n        criteria_preds_wide_df[criteria_preds_wide_df['Model'] == model_key]\n    )).rename_axis('metric').reset_index(level = 'metric')\n    model_scores_df['Model'] = model_key\n    refined_criteria_score_parts.append(model_scores_df)\nrefined_criteria_scores_df = (\n    pd.concat(refined_criteria_score_parts)\n        .melt(id_vars = ['Model', 'metric'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['Model', 'Finding'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\nrefined_criteria_scores_df['support_pct'] = refined_criteria_scores_df['support'] / refined_baseline_scores_df['support']\nrefined_criteria_scores_df['Condition'] = 'criteria-only-labeled'\n\nrefined_masked_score_parts = []\nfor model_key in REFINED_MODEL_KEYS:\n    model_scores_df = pd.DataFrame(get_scores_by_class(\n        gold_df,\n        refined_preds_wide_df[refined_preds_wide_df['Model'] == model_key]\n    )).rename_axis('metric').reset_index(level = 'metric')\n    model_scores_df['Model'] = model_key\n    refined_masked_score_parts.append(model_scores_df)\nrefined_masked_scores_df = (\n    pd.concat(refined_masked_score_parts)\n        .melt(id_vars = ['Model', 'metric'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['Model', 'Finding'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\nrefined_masked_scores_df['support_pct'] = refined_masked_scores_df['support'] / refined_baseline_scores_df['support']\nrefined_masked_scores_df['Condition'] = 'refined-only-labeled'\n\nrefined_scores_df = pd.concat([refined_baseline_scores_df, refined_criteria_scores_df, refined_masked_scores_df])\nrefined_scores_df[['support', 'total_labeled']] = refined_scores_df[['support', 'total_labeled']].astype('Int64')\nrefined_scores_df['Condition'] = pd.Categorical(refined_scores_df['Condition'], categories = ['baseline', 'criteria-only-labeled', 'refined-only-labeled'], ordered = True)\n\nREFINED_CONDITION_COLORS = {'baseline': '#878f99', 'criteria-only-labeled': '#daa3e9', 'refined-only-labeled': '#1f77b4'}\nREFINED_CONDITION_LABELS = ['Baseline', 'Criteria - Abstains Masked', 'Refined - Abstains Masked']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:51.224646Z","iopub.execute_input":"2026-09-17T20:14:51.224957Z","iopub.status.idle":"2026-09-17T20:14:53.858275Z","shell.execute_reply.started":"2026-09-17T20:14:51.224918Z","shell.execute_reply":"2026-09-17T20:14:53.857416Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(refined_scores_df, aes(x = 'Model', y = 'support_pct', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + scale_x_discrete(limits = refined_model_order_by_abstention)\n        + scale_fill_manual(\n            values = REFINED_CONDITION_COLORS,\n            name=\"Support\",\n            labels = REFINED_CONDITION_LABELS\n        )\n        + labs(x=\"LLM Model\", y=\"Proportion\", title=\"Support (Positive Labels) Before and After Abstention Masking\", fill = \"Support\")\n        + facet_wrap('Finding')\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(10, 10))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:53.859334Z","iopub.execute_input":"2026-09-17T20:14:53.859604Z","iopub.status.idle":"2026-09-17T20:14:55.096567Z","shell.execute_reply.started":"2026-09-17T20:14:53.859577Z","shell.execute_reply":"2026-09-17T20:14:55.095577Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(\n        (\n            refined_scores_df[['Condition', 'Model', 'f1-score', 'precision', 'recall', 'balanced_accuracy_score', 'npv', 'specificity']]\n                .groupby(['Condition', 'Model'])\n                .mean()\n                .reset_index()\n                .melt(id_vars = ['Condition', 'Model'], var_name = 'Metric', value_name = 'Value')\n        ),\n        aes(x = 'Model', y = 'Value', fill = 'Condition')\n    )\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Metric', ncol = 3)\n        + scale_x_discrete(limits = refined_model_order_by_abstention)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = REFINED_CONDITION_COLORS,\n            name=\"Condition\",\n            labels = REFINED_CONDITION_LABELS\n        )\n        + labs(x=\"LLM Model\", y=\"Macro Average\", title=\"Overall Performance by Model and Metric - Refined Program\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(10, 5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:55.097633Z","iopub.execute_input":"2026-09-17T20:14:55.097904Z","iopub.status.idle":"2026-09-17T20:14:55.843099Z","shell.execute_reply.started":"2026-09-17T20:14:55.097877Z","shell.execute_reply":"2026-09-17T20:14:55.842126Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(refined_scores_df, aes(x = 'Model', y = 'precision', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = refined_model_order_by_abstention)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = REFINED_CONDITION_COLORS,\n            name=\"Condition\",\n            labels = REFINED_CONDITION_LABELS\n        )\n        + labs(x=\"LLM Model\", y=\"Precision\", title=\"Positive Predictive Value (Precision) by Condition, Model, and Class\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:55.844265Z","iopub.execute_input":"2026-09-17T20:14:55.844615Z","iopub.status.idle":"2026-09-17T20:14:57.090294Z","shell.execute_reply.started":"2026-09-17T20:14:55.844573Z","shell.execute_reply":"2026-09-17T20:14:57.08931Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(refined_scores_df, aes(x = 'Model', y = 'npv', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = refined_model_order_by_abstention)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = REFINED_CONDITION_COLORS,\n            name=\"Condition\",\n            labels = REFINED_CONDITION_LABELS\n        )\n        + labs(x=\"LLM Model\", y=\"NPV\", title=\"Negative Predictive Value (NPV) by Condition, Model, and Class\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:57.09141Z","iopub.execute_input":"2026-09-17T20:14:57.091733Z","iopub.status.idle":"2026-09-17T20:14:58.319585Z","shell.execute_reply.started":"2026-09-17T20:14:57.091704Z","shell.execute_reply":"2026-09-17T20:14:58.31862Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- No obvious gains here with the program and any differences between model are more than likely to be due noise than anything else\n- The five-category evidence basis Enum is too complex to generate data that could be used a feature for a linear classifier; similarly\n  abstentions are too infrequent to be used in a similar way","metadata":{}},{"cell_type":"markdown","source":"## 11. Consensus Labeling Program\n- Consolidate the three prior programs into a single program for pseudolabeling the unlabeled studies\n- Class definitions are the criteria set, carried forward unchanged; the Section 9 revisions do not\n  carry over\n- The annotation enum is two-way: abstention generally seems to produce noise without value, so `UNDETERMINED` is removed\n  from the schema rather than discouraged in the instructions\n- `evidence_basis` carries forward from the refined program, simplified to three categories -\n  `NOT_EVALUABLE` existed only to license abstention, and `THRESHOLD` folds into `STATED` and\n  `DESCRIBED`, with the severity reasoning left to the rationale\n- Confidence takes the refined program's definition: confidence in the annotation under the\n  annotation policy, not confidence that the finding is present\n- Self-contained: every type is redeclared here rather than reused from the earlier programs for ease","metadata":{}},{"cell_type":"code","source":"class ClassificationLabelConsensus(Enum):\n    PRESENT = 1\n    ABSENT = 0\n\n\nclass ConfidenceRatingConsensus(Enum):\n    VERY_HIGH = 5\n    HIGH = 4\n    MODERATE = 3\n    LOW = 2\n    VERY_LOW = 1\n    NO_CONFIDENCE = 0\n\n\nclass EvidenceBasisConsensus(Enum):\n    \"\"\"The kind of report evidence the annotation rests on. Independent of the\n    annotation: STATED and DESCRIBED can each support PRESENT or ABSENT.\"\"\"\n\n    STATED = \"the report names the finding and either asserts or denies it\"\n    DESCRIBED = \"the finding is not named; the annotation follows from described morphology, or from the relevant structure being described as normal\"\n    NOT_MENTIONED = \"the report does not mention this finding\"\n\n\nclass FindingAnnotationConsensus(TypedDict):\n    annotation: Annotated[\n        ClassificationLabelConsensus,\n        Field(\n            description=\n\"\"\"PRESENT or ABSENT for this finding, per the finding definition. Every class takes one of the\ntwo; there is no abstention.\"\"\"\n        ),\n    ]\n    confidence: Annotated[\n        ConfidenceRatingConsensus,\n        Field(\n            description=\n\"\"\"Confidence in the specific annotation made based on the annotation policy\"\"\"\n        ),\n    ]\n    evidence_basis: Annotated[\n        EvidenceBasisConsensus,\n        Field(\n            description=\n\"\"\"The kind of report evidence this annotation rests on. Select the most specific applicable\ncategory: prefer STATED or DESCRIBED, and use NOT_MENTIONED only when neither applies.\"\"\"\n        ),\n    ]\n    rationale: Annotated[\n        str,\n        Field(\n            description=\n\"\"\"Short justification (in English, regardless of the report language) for the assigned\nannotation and confidence. Quote or closely paraphrase the report language relied on, or state\nthat the report is silent.\"\"\"\n        ),\n    ]\n\n\nclass ReportFindingsConsensus(TypedDict):\n    ACL: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A high-grade partial or full-thickness tear of the anterior cruciate ligament, meaning complete discontinuity\nof the ligament, or more than 50 percent of fibers disrupted, with or without secondary signs such as characteristic\npivot-shift bone contusions. Mild signal change, degeneration, or thickening without discontinuity is graded negative.\"\"\"\n        ),\n    ]\n    MCL: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A high-grade partial or complete acute tear of the medial collateral ligament, with disrupted fibers\nand edema within and adjacent to the ligament. Low-grade sprains and chronic or remote stress changes are graded negative.\"\"\"\n        ),\n    ]\n    Medial_Meniscus: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the medial meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Lateral_Meniscus: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"Abnormal signal that definitely contacts the meniscal surface on at least two images, or a morphologic abnormality such as a\ntruncated, diminutive, or displaced fragment, involving the lateral meniscus. Intrasubstance degeneration that does not reach the\nsurface is negative.\"\"\"\n        ),\n    ]\n    Medial_OA: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the medial compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Lateral_OA: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the lateral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    PF_OA: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large area (roughly 1 cm or greater) of high-grade cartilage loss, defined as greater than 50 percent\nof cartilage thickness, in the patellofemoral compartment, with or without underlying subchondral marrow changes.\"\"\"\n        ),\n    ]\n    Effusion: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large amount of fluid distending the joint.\"\"\"\n        ),\n    ]\n    Synovitis: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"Inflammation and thickening of the synovial lining of the joint.\"\"\"\n        ),\n    ]\n    Baker: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A moderate or large fluid collection in the characteristic Baker (popliteal) cyst location behind the knee.\"\"\"\n        ),\n    ]\n    Contusion: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"A bone contusion, seen as bone marrow edema-like signal from impact, without a discrete fracture line.\"\"\"\n        ),\n    ]\n    Fracture: Annotated[\n        FindingAnnotationConsensus,\n        Field(\n            description=\n\"\"\"An acute cortical break or fracture line.\"\"\"\n        ),\n    ]\n\n\n# Same 12 names as FINDING_COLUMNS, derived from the schema so the two cannot drift apart\nFINDING_NAMES_CONSENSUS = tuple(ReportFindingsConsensus.__annotations__)\nENUM_KEYS_CONSENSUS = {'annotation': ClassificationLabelConsensus, 'confidence': ConfidenceRatingConsensus}\n\n\ndef _as_int_consensus(value, enum_cls: type[Enum]) -> int | None:\n    if value is None:\n        return None\n    if isinstance(value, enum_cls):\n        return int(value.value)\n    if isinstance(value, str):\n        name = value.strip().upper()\n        if name in enum_cls.__members__:\n            return int(enum_cls[name].value)\n        value = int(value)\n    return int(enum_cls(int(value)).value)\n\n\ndef _as_evidence_basis_name_consensus(value) -> str | None:\n    if value is None:\n        return None\n    if isinstance(value, EvidenceBasisConsensus):\n        return value.name\n    name = str(value).strip().upper()\n    if name in EvidenceBasisConsensus.__members__:\n        return name\n    return EvidenceBasisConsensus(str(value)).name\n\n\ndef as_long_records_consensus(findings: dict, study_instance_uid: str) -> list[dict]:\n    records = []\n    for name in FINDING_NAMES_CONSENSUS:\n        annotation = findings.get(name) or {}\n        rationale = annotation.get('rationale')\n        records.append({\n            'StudyInstanceUID': str(study_instance_uid),\n            'Finding': name,\n            **{key: _as_int_consensus(annotation.get(key), cls) for key, cls in ENUM_KEYS_CONSENSUS.items()},\n            'evidence_basis': _as_evidence_basis_name_consensus(annotation.get('evidence_basis')),\n            'rationale': None if rationale is None else str(rationale),\n        })\n    return records\n\n\nclass KneeFindingsConsensus(BaseModel):\n    findings: ReportFindingsConsensus = Field(\n        description=\n\"\"\"Annotations per class for the knee MRI radiology report\"\"\"\n    )\n\n\nclass KneeMRIOutputConsensus(dspy.Signature):\n    \"\"\"\n    Act as an expert musculoskeletal radiologist annotating a knee MRI radiology report for\n    each class of finding. Use the definition supplied with each class when making annotations.\n\n    DECISION PROCEDURE\n    For each class determine in order whether the report text provides:\n      1. A clear and direct statement confirming the presence of the finding (or its absence)\n         according to the finding criteria, including any severity thresholds?\n      2. Equivalent language describing the finding's presence (or absence) (eg, using clinically\n         related language or as part of a composite diagnosis)?\n      3. A normal or negative statement about the relevant structure or compartment that would\n         indicate the finding is absent (eg, the relevant structure is described as normal or\n         without defects)?\n      4. Another reported finding that makes this one likely or unlikely, given normal co-occurrence\n         of knee pathology?\n      5. Silence on the finding? If none of the above applies, annotate ABSENT.\n    Give the impression/conclusion precedence over the body when they conflict, and say so in\n    the rationale.\n\n    ANNOTATION POLICY\n    Every class receives PRESENT or ABSENT; there is no abstention. A report that is silent,\n    equivocal, or non-diagnostic for a class still receives the more likely of the two labels, with\n    the residual uncertainty carried by the confidence rating rather than by withholding a label.\n    - PRESENT: the report affirms the finding, describes its substance, or supports it strongly\n      enough that an independent image reader would more likely than not have recorded it.\n    - ABSENT: the report denies the finding, describes the relevant structures as normal, or is\n      silent about it.\n\n    HOW TO HANDLE SILENCE\n    - Silence resolves to ABSENT.\n    - Silence is weaker evidence of absence for findings that are frequently observed but omitted as\n      incidental (synovial change, compartment-specific cartilage grading in a trauma report).\n      Annotate ABSENT in those cases as well, and carry the weakness in a lower confidence rating.\n\n    HOW TO HANDLE SEVERITY QUALIFIERS\n      - Where a class definition has a size or severity threshold, do not resolve to ABSENT on the\n        strength of a diminutive adjective alone when the finding itself is affirmed. Weigh the\n        objective content of the description (measurements, extent, number of regions, secondary\n        signs) more importantly.\n      - Distinguish a diminutive qualifier from a negation. \"Mild effusion\" affirms the finding;\n        \"no significant effusion\" and \"no effusion\" indicate the finding is ABSENT.\n      - A finding that is named but not quantified should generally be graded PRESENT, as\n        radiologists tend to name findings that are conspicuous.\n      - The exception is the cartilage/osteoarthritis classes, where the grading vocabulary is\n        standardized (Outerbridge/ICRS grade, \"full-thickness\", \"advanced\", \"complete cartilage\n        loss\") and does track the reference threshold. There, respect the reported grade, and\n        treat low-grade chondral change, fissuring, thinning, an unqualified diagnosis of\n        \"osteoarthritis\", or isolated osteophytes as not meeting the definition.\n\n    FINDINGS ARE NOT MUTUALLY EXCLUSIVE\n    Annotate each class independently against its own definition. In particular, a report that\n    attributes an abnormality to one entity does not thereby exclude an overlapping class: bone\n    marrow edema accompanying a fracture still satisfies a marrow-edema definition, and\n    traumatic and degenerative explanations for the same signal can both be recorded.\n\n    CONFIDENCE\n    Confidence is the confidence in the specific annotation according to the policy above; it is\n    NOT confidence that a finding is present. Annotations of PRESENT and of ABSENT will be made\n    with varying levels of confidence, which is what the rating should reflect. An ABSENT\n    annotation resting on silence alone, or a call that turns on an unquantified severity\n    qualifier, belongs at the low end of the scale.\n\n    EVIDENCE BASIS\n    Commit to the kind of report evidence the annotation rests on before writing the rationale.\n    STATED and DESCRIBED each support PRESENT or ABSENT; NOT_MENTIONED records that the report is\n    silent, and pairs with an ABSENT annotation under the silence policy.\n\n    RATIONALE\n    The rationale must quote or closely paraphrase the specific report language relied on, or state\n    that the report is silent. It must be consistent with the evidence basis and with the annotation.\n    \"\"\"\n    report: str = dspy.InputField()\n    classification: KneeFindingsConsensus = dspy.OutputField()\n\n\nclass ConsensusClassifier(dspy.Module):\n    def __init__(self):\n        super().__init__()\n        self.classification_module = dspy.ChainOfThought(KneeMRIOutputConsensus)\n\n    def forward(self, report: str, study_instance_uid: str) -> dspy.Prediction:\n        result = self.classification_module(report=report)\n\n        return dspy.Prediction(\n            study_instance_uid=study_instance_uid,\n            findings=as_long_records_consensus(result.classification.findings, study_instance_uid),\n        )\n\ndef format_examples_consensus(df: pd.DataFrame) -> list[dspy.Example]:\n    df_findings = df[FINDING_COLUMNS].astype(int).to_dict(orient = 'records')\n    examples = []\n    for i, row in enumerate(df.itertuples()):\n        example = dspy.Example(\n            study_instance_uid = row.StudyInstanceUID,\n            report = row.Report,\n            findings = df_findings[i]\n        ).with_inputs(\"report\", \"study_instance_uid\")\n        examples.append(example)\n    return examples","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:58.321058Z","iopub.execute_input":"2026-09-17T20:14:58.321388Z","iopub.status.idle":"2026-09-17T20:14:58.354055Z","shell.execute_reply.started":"2026-09-17T20:14:58.321358Z","shell.execute_reply":"2026-09-17T20:14:58.353164Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Runs predictions using the consensus prompt across the three carried-forward models, unless the\n# serialized output is available\n\nconsensus_preds_load_path = DATA_LOC / 'datasets' / 'joshuaziel' / 'rsna-knee-mri-radiology-report-classifications' / 'raw_prediction_results_v4.json'\nconsensus_preds_persist_path = OUTPUTS_LOC / 'raw_prediction_results_v4.json'\n\nCONSENSUS_MODEL_KEYS = REFINED_MODEL_KEYS\n\nif consensus_preds_load_path.is_file():\n    with open(str(consensus_preds_load_path), 'r') as f:\n        consensus_preds_raw = json.load(f)\n    print(\"Loaded LM predictions from file\")\n\n# Same hand-rolled retry loop as the earlier runs\nelse:\n    program = ConsensusClassifier()\n    consensus_preds_raw = {}\n    for model_key in tqdm(CONSENSUS_MODEL_KEYS, dynamic_ncols = True):\n        model_spec = MODELS[model_key]\n        dspy_lm = dspy.LM(model = model_spec.name, temperature = 1.0, model_type = model_spec.type, api_key = model_spec.key, cache = False)\n        dspy.configure(lm=dspy_lm)\n        model_preds = []\n\n        for row in gold_df.itertuples():\n            success, tries = False, 0\n            while not success and tries < 4:\n                tries += 1\n                try:\n                    prediction = program(study_instance_uid=row.StudyInstanceUID, report=row.Report)\n                    model_preds.extend(prediction.findings)\n                    success = True\n                except Exception as e:\n                    tqdm.write(f'Error predicting labels for {row.StudyInstanceUID}: {e}')\n            if not success:\n                model_preds.extend(as_long_records_consensus({}, row.StudyInstanceUID))\n        consensus_preds_raw[model_key] = model_preds\n\n    with open(str(consensus_preds_persist_path), 'w') as f:\n        json.dump(consensus_preds_raw, f, indent = 4)\n\nconsensus_preds_long_df = pd.concat(\n    [pd.DataFrame(records).assign(Model=model_key) for model_key, records in consensus_preds_raw.items()],\n    ignore_index=True,\n)\nconsensus_preds_long_df = consensus_preds_long_df[consensus_preds_long_df['StudyInstanceUID'].isin(gold_train_ids['StudyInstanceUID'])]\nconsensus_preds_long_df['annotation'] = consensus_preds_long_df['annotation'].astype(\"Int64\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:58.355175Z","iopub.execute_input":"2026-09-17T20:14:58.355545Z","iopub.status.idle":"2026-09-17T20:14:58.409851Z","shell.execute_reply.started":"2026-09-17T20:14:58.355514Z","shell.execute_reply":"2026-09-17T20:14:58.408974Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 12. Evaluation of the Consensus Program\n- Does forcing a two-way call change agreement between the three models relative to the refined\n  program, which let them abstain?\n- Where does confidence land once uncertainty has nowhere to go but the confidence rating?\n- What evidence does the consensus program rest its annotations on, with the threshold and\n  non-evaluable categories removed?\n- How does performance compare to baseline and to the two prior programs with their abstains masked?\n- Evaluation is restricted to the train ids, as in Sections 6 and 10","metadata":{}},{"cell_type":"code","source":"make_cohens_kappa_matrix_plot(\n    consensus_preds_long_df, # no abstentions to fill, so the long frame is passed directly\n    models = {model_key: MODELS[model_key] for model_key in CONSENSUS_MODEL_KEYS},\n    mid_value = 0.7, # 0.7 might generally be consider to be a moderate level of agreement\n    scale_limits = (0.00, 1.0)\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:58.411105Z","iopub.execute_input":"2026-09-17T20:14:58.41181Z","iopub.status.idle":"2026-09-17T20:14:58.747105Z","shell.execute_reply.started":"2026-09-17T20:14:58.411778Z","shell.execute_reply":"2026-09-17T20:14:58.745996Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(consensus_preds_long_df, aes(x=\"confidence\"))\n        + geom_histogram(bins = 6, color = \"white\", fill = \"#1f77b4\")\n        + facet_grid(cols = 'Finding', rows = 'Model')\n        + scale_x_continuous(breaks=[1,3,5], labels=[\"Very Low\", \"Moderate\", \"Very High\"])\n        + theme(figure_size=(15, 4))\n        + labs(x=\"Confidence Level\", y=\"Count, n\", title=\"Frequency Distribution of Confidence Levels by LLM and Class - Consensus Program\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:14:58.748277Z","iopub.execute_input":"2026-09-17T20:14:58.748636Z","iopub.status.idle":"2026-09-17T20:15:02.432296Z","shell.execute_reply.started":"2026-09-17T20:14:58.748605Z","shell.execute_reply":"2026-09-17T20:15:02.431314Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"CONSENSUS_EVIDENCE_BASIS_ORDER = [evidence_basis.name for evidence_basis in EvidenceBasisConsensus]\nCONSENSUS_EVIDENCE_BASIS_COLORS = {\n    'STATED': '#1f77b4',\n    'DESCRIBED': '#6baed6',\n    'NOT_MENTIONED': '#878f99',\n}\n(\n    ggplot(\n        consensus_preds_long_df.groupby(['Model', 'Finding', 'evidence_basis']).size().reset_index(name = 'n'),\n        aes(x = 'Model', y = 'n', fill = 'evidence_basis')\n    )\n    + facet_wrap('Finding')\n    + geom_col(position = 'fill', color = 'white')\n    + scale_x_discrete(limits = CONSENSUS_MODEL_KEYS)\n    + scale_fill_manual(values = CONSENSUS_EVIDENCE_BASIS_COLORS, limits = CONSENSUS_EVIDENCE_BASIS_ORDER, name = \"Evidence Basis\")\n    + labs(x=\"LLM Model\", y=\"Proportion of Annotations\", title=\"Evidence Basis by Class and LLM - Consensus Program\")\n    + theme(axis_text_x=element_text(angle=45, hjust=1),\n        figure_size=(10, 8))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:02.433474Z","iopub.execute_input":"2026-09-17T20:15:02.43376Z","iopub.status.idle":"2026-09-17T20:15:03.852795Z","shell.execute_reply.started":"2026-09-17T20:15:02.43373Z","shell.execute_reply":"2026-09-17T20:15:03.851412Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"consensus_preds_wide_df = consensus_preds_long_df.pivot(index = ['Model', 'StudyInstanceUID'], columns = 'Finding', values = 'annotation').reset_index()\n\nconsensus_score_parts = []\nfor model_key in CONSENSUS_MODEL_KEYS:\n    model_scores_df = pd.DataFrame(get_scores_by_class(\n        gold_df,\n        consensus_preds_wide_df[consensus_preds_wide_df['Model'] == model_key]\n    )).rename_axis('metric').reset_index(level = 'metric')\n    model_scores_df['Model'] = model_key\n    consensus_score_parts.append(model_scores_df)\nconsensus_only_scores_df = (\n    pd.concat(consensus_score_parts)\n        .melt(id_vars = ['Model', 'metric'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['Model', 'Finding'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\nconsensus_only_scores_df['support_pct'] = consensus_only_scores_df['support'] / refined_baseline_scores_df['support']\nconsensus_only_scores_df['Condition'] = 'consensus'\n\n# The three comparison conditions are carried over from Section 10 rather than recomputed\nconsensus_scores_df = pd.concat([\n    refined_baseline_scores_df,\n    refined_criteria_scores_df,\n    refined_masked_scores_df,\n    consensus_only_scores_df,\n])\nconsensus_scores_df[['support', 'total_labeled']] = consensus_scores_df[['support', 'total_labeled']].astype('Int64')\nconsensus_scores_df['Condition'] = pd.Categorical(\n    consensus_scores_df['Condition'],\n    categories = ['baseline', 'criteria-only-labeled', 'refined-only-labeled', 'consensus'],\n    ordered = True\n)\n\nCONSENSUS_CONDITION_COLORS = {\n    'baseline': '#878f99',\n    'criteria-only-labeled': '#daa3e9',\n    'refined-only-labeled': '#6baed6',\n    'consensus': '#1f77b4',\n}\nCONSENSUS_CONDITION_LABELS = ['Baseline', 'Criteria - Abstains Masked', 'Refined - Abstains Masked', 'Consensus']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:03.859252Z","iopub.execute_input":"2026-09-17T20:15:03.859618Z","iopub.status.idle":"2026-09-17T20:15:04.829916Z","shell.execute_reply.started":"2026-09-17T20:15:03.859585Z","shell.execute_reply":"2026-09-17T20:15:04.828773Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(\n        (\n            consensus_scores_df[['Condition', 'Model', 'f1-score', 'precision', 'recall', 'balanced_accuracy_score', 'npv', 'specificity']]\n                .groupby(['Condition', 'Model'])\n                .mean()\n                .reset_index()\n                .melt(id_vars = ['Condition', 'Model'], var_name = 'Metric', value_name = 'Value')\n        ),\n        aes(x = 'Model', y = 'Value', fill = 'Condition')\n    )\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Metric', ncol = 3)\n        + scale_x_discrete(limits = CONSENSUS_MODEL_KEYS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = CONSENSUS_CONDITION_COLORS,\n            name=\"Condition\",\n            labels = CONSENSUS_CONDITION_LABELS\n        )\n        + labs(x=\"LLM Model\", y=\"Macro Average\", title=\"Overall Performance by Model and Metric - Consensus Program\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(10, 5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:04.831171Z","iopub.execute_input":"2026-09-17T20:15:04.831492Z","iopub.status.idle":"2026-09-17T20:15:05.613587Z","shell.execute_reply.started":"2026-09-17T20:15:04.831461Z","shell.execute_reply":"2026-09-17T20:15:05.612402Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(consensus_scores_df, aes(x = 'Model', y = 'precision', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = CONSENSUS_MODEL_KEYS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = CONSENSUS_CONDITION_COLORS,\n            name=\"Condition\",\n            labels = CONSENSUS_CONDITION_LABELS\n        )\n        + labs(x=\"LLM Model\", y=\"Precision\", title=\"Positive Predictive Value (Precision) by Condition, Model, and Class\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:05.614951Z","iopub.execute_input":"2026-09-17T20:15:05.615328Z","iopub.status.idle":"2026-09-17T20:15:06.898272Z","shell.execute_reply.started":"2026-09-17T20:15:05.615287Z","shell.execute_reply":"2026-09-17T20:15:06.897355Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(consensus_scores_df, aes(x = 'Model', y = 'npv', fill = 'Condition'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = CONSENSUS_MODEL_KEYS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = CONSENSUS_CONDITION_COLORS,\n            name=\"Condition\",\n            labels = CONSENSUS_CONDITION_LABELS\n        )\n        + labs(x=\"LLM Model\", y=\"NPV\", title=\"Negative Predictive Value (NPV) by Condition, Model, and Class\", fill = \"Condition\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:06.89955Z","iopub.execute_input":"2026-09-17T20:15:06.899977Z","iopub.status.idle":"2026-09-17T20:15:08.180838Z","shell.execute_reply.started":"2026-09-17T20:15:06.899929Z","shell.execute_reply":"2026-09-17T20:15:08.179876Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Takeaways\n- Not unexpectedly at this point, a pretty big nothing-burger.  The only reason to use the consensus program is\n  - It avoids adding in a lot of additional content into the instructions that add little value\n  - It provides features that could be useful in a linear classifier\n- Generally, any numeric differences by model are not substantial enough to warrant the signficiantly higher cost\n  of terra or sonnet - took a while to get here...probably too long\n\n## Next Steps\n- Without a model/mapping function to set probabilities, macro avg BA is essentially the competition metric and its literally flat across models\n  and conditions\n- I'm pretty set that the model for predictions will be 'luna' as a practical matter given cost and similarity in performance; the only reason to \n  use one dspy program over another is if it provides features that could help performance as part of, for example a linear classifier\n  - My gut says the two options are the criteria model or the refined model\n  - The 'criteria' program because we're going to label 4K+ records and I think the LLM should have the actual criteria availabel...even though I acknowledge there\n    is little reason to think this...call me superstitious\n  - The 'refined' model because it offers perhaps the richest features that could be used for such a linear classifier.","metadata":{}},{"cell_type":"markdown","source":"## 13. Logistic Regression Classification for Criteria and Consensus Conditions\n- Leverage the available features and create simple classification models for evaluation","metadata":{}},{"cell_type":"code","source":"#Modeling Helpers\n\ndef add_criteria_features(modeling_df: pd.DataFrame) -> pd.DataFrame:\n    return modeling_df.assign(\n        anno_present = lambda x: x['annotation'].apply(lambda a: 0 if pd.isna(a) else a).astype('Int64'),\n        anno_abstained = lambda x: x['annotation'].apply(lambda a: 1 if pd.isna(a) else 0).astype('Int64')\n    )\n\n\ndef add_consensus_features(modeling_df: pd.DataFrame) -> pd.DataFrame:\n    return modeling_df.assign(\n        anno_duped = lambda x: x['annotation'],\n        lo_confidence = lambda x: x['confidence'].apply(lambda c: 0 if c == 5 else 1).astype('Int64'),\n        lo_evidence = lambda x: x['evidence_basis'].apply(lambda e: 0 if e == 'STATED' else 1).astype('Int64')\n    )\n\n\n# Each condition differs only in which prediction fields survive into the pivot index and which\n# features it derives from them.  The dummy-encoded findings are common to both.\nCONDITION_SPECS = {\n    'criteria': {\n        'index_cols': ['StudyInstanceUID', 'Finding', 'annotation', 'gold_label'],\n        'extra_features': ['anno_present', 'anno_abstained'],\n        'add_features': add_criteria_features\n    },\n    'consensus': {\n        'index_cols': ['StudyInstanceUID', 'Finding', 'annotation', 'gold_label', 'confidence', 'evidence_basis'],\n        'extra_features': ['anno_duped', 'lo_confidence', 'lo_evidence'],\n        'add_features': add_consensus_features\n    }\n}\n\nRESULT_COLUMNS = ['StudyInstanceUID', 'Finding', 'annotation', 'model_labels', 'model_probs',\n                  'gold_label', 'condition', 'set']\n\n\ndef build_gold_long(gold_df: pd.DataFrame, study_ids: pd.Series) -> pd.DataFrame:\n    return (gold_df[gold_df['StudyInstanceUID'].isin(study_ids)]\n        .drop(columns = ['Report', 'Is_English', 'Num_Findings', 'Language', 'Above_Median_Findings'])\n        .melt(id_vars = 'StudyInstanceUID', var_name = 'Finding', value_name = 'gold_label')\n        .assign(gold_label = lambda x: x['gold_label'].astype('Int64'))\n    )\n\n\ndef build_preds_long(preds_raw: dict, study_ids: pd.Series, model_key: str = 'luna') -> pd.DataFrame:\n    preds_df = pd.concat(\n        [pd.DataFrame(records).assign(Model = key) for key, records in preds_raw.items()],\n        ignore_index = True\n    )\n    return (preds_df[(preds_df['Model'] == model_key) & (preds_df['StudyInstanceUID'].isin(study_ids))]\n        .drop(columns = ['Model', 'rationale'])\n        .assign(annotation = lambda x: x['annotation'].astype('Int64'))\n    )\n\n\ndef build_modeling_frame(preds_long_df: pd.DataFrame, gold_long_df: pd.DataFrame, condition: str) -> pd.DataFrame:\n    spec = CONDITION_SPECS[condition]\n    return (pd.merge(\n            preds_long_df,\n            gold_long_df,\n            on = ['StudyInstanceUID', 'Finding'],\n            how = 'left'\n        )\n        .assign(\n            finding_dupe = lambda x: x['Finding'],\n            dummy = 1\n        )\n        .pivot(index = spec['index_cols'], columns = 'finding_dupe', values = 'dummy')\n        .fillna(0)\n        .astype('Int64')\n        .reset_index()\n        .pipe(spec['add_features'])\n    )\n\n\ndef apply_model(modeling_df: pd.DataFrame, condition: str, model: LogisticRegression | None = None) -> tuple:\n    features = modeling_df[FINDING_COLUMNS + CONDITION_SPECS[condition]['extra_features']].to_numpy()\n    if model is None:\n        model = LogisticRegression().fit(features, modeling_df['gold_label'].to_numpy())\n    scored_df = modeling_df.assign(\n        model_labels = model.predict(features),\n        model_probs = model.predict_proba(features)[:, 1]\n    )\n    return scored_df, model\n#Train and Test Set Predictions\n\ngold_long_train_df = build_gold_long(gold_df, gold_train_ids['StudyInstanceUID'])\ngold_long_test_df = build_gold_long(gold_df, gold_test_ids['StudyInstanceUID'])\n\nluna_frames = {}\nluna_models = {}\nfor condition, preds_raw in [('criteria', criteria_preds_raw), ('consensus', consensus_preds_raw)]:\n    train_frame = build_modeling_frame(\n        build_preds_long(preds_raw, gold_train_ids['StudyInstanceUID']),\n        gold_long_train_df,\n        condition\n    )\n    test_frame = build_modeling_frame(\n        build_preds_long(preds_raw, gold_test_ids['StudyInstanceUID']),\n        gold_long_test_df,\n        condition\n    )\n    # The test set is scored with the model fit on the train set, never refit\n    train_frame, luna_models[condition] = apply_model(train_frame, condition)\n    test_frame, _ = apply_model(test_frame, condition, luna_models[condition])\n    luna_frames[condition, 'train'] = train_frame.assign(condition = condition, set = 'train')\n    luna_frames[condition, 'test'] = test_frame.assign(condition = condition, set = 'test')\n\ntrain_test_model_results = pd.concat(\n    [frame[RESULT_COLUMNS] for frame in luna_frames.values()],\n    ignore_index = True\n)\ntrain_test_model_gold = (\n    train_test_model_results[['StudyInstanceUID', 'Finding', 'gold_label', 'condition', 'set']]\n        .pivot(index = ['StudyInstanceUID', 'condition', 'set'], columns = 'Finding', values = 'gold_label')\n        .reset_index()\n)\n\ntrain_test_model_probs = (\n    train_test_model_results[['StudyInstanceUID', 'Finding', 'model_probs', 'condition', 'set']]\n        .pivot(index = ['StudyInstanceUID', 'condition', 'set'], columns = 'Finding', values = 'model_probs')\n        .reset_index()\n)\n\ntrain_test_model_labels = (\n    train_test_model_results[['StudyInstanceUID', 'Finding', 'model_labels', 'condition', 'set']]\n        .pivot(index = ['StudyInstanceUID', 'condition', 'set'], columns = 'Finding', values = 'model_labels')\n        .astype('Int64')\n        .reset_index()\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:08.182115Z","iopub.execute_input":"2026-09-17T20:15:08.182511Z","iopub.status.idle":"2026-09-17T20:15:08.437871Z","shell.execute_reply.started":"2026-09-17T20:15:08.182472Z","shell.execute_reply":"2026-09-17T20:15:08.434641Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#Label Model Scores by Condition and Set\n\nTRAIN_TEST_SET_ORDER = ['train', 'test']\nTRAIN_TEST_SET_COLORS = {\n    'train': '#1f77b4',\n    'test': '#878f99',\n}\nCONDITION_ORDER = list(CONDITION_SPECS)\nCONDITION_LABELS = ['Criteria', 'Consensus']\nMACRO_AVERAGE_LABEL = 'Macro Average'\n\ntrain_test_score_parts = []\nfor condition in CONDITION_ORDER:\n    for set_name in TRAIN_TEST_SET_ORDER:\n        set_scores_df = pd.DataFrame(get_scores_by_class(\n            train_test_model_gold[(train_test_model_gold['condition'] == condition) & (train_test_model_gold['set'] == set_name)],\n            train_test_model_labels[(train_test_model_labels['condition'] == condition) & (train_test_model_labels['set'] == set_name)]\n        )).rename_axis('metric').reset_index(level = 'metric')\n        set_scores_df['condition'] = condition\n        set_scores_df['set'] = set_name\n        train_test_score_parts.append(set_scores_df)\n\ntrain_test_scores_df = (\n    pd.concat(train_test_score_parts)\n        .melt(id_vars = ['condition', 'set', 'metric'], var_name = 'Finding', value_name = 'Value')\n        .pivot(index = ['condition', 'set', 'Finding'], columns = 'metric', values = 'Value')\n        .reset_index()\n)\ntrain_test_scores_df[['support', 'total_labeled']] = train_test_scores_df[['support', 'total_labeled']].astype('Int64')\ntrain_test_scores_df['set'] = pd.Categorical(train_test_scores_df['set'], categories = TRAIN_TEST_SET_ORDER, ordered = True)\ntrain_test_scores_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:08.439391Z","iopub.execute_input":"2026-09-17T20:15:08.439803Z","iopub.status.idle":"2026-09-17T20:15:09.722097Z","shell.execute_reply.started":"2026-09-17T20:15:08.439765Z","shell.execute_reply":"2026-09-17T20:15:09.720962Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(\n        (\n            train_test_scores_df[['condition', 'set', 'f1-score', 'precision', 'recall', 'balanced_accuracy_score', 'npv', 'specificity']]\n                .groupby(['condition', 'set'], observed = True)\n                .mean()\n                .reset_index()\n                .melt(id_vars = ['condition', 'set'], var_name = 'Metric', value_name = 'Value')\n        ),\n        aes(x = 'condition', y = 'Value', fill = 'set')\n    )\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Metric', ncol = 3)\n        + scale_x_discrete(limits = CONDITION_ORDER, labels = CONDITION_LABELS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = TRAIN_TEST_SET_COLORS,\n            limits = TRAIN_TEST_SET_ORDER,\n            name = \"Set\",\n            labels = ['Train', 'Test']\n        )\n        + labs(x=\"Condition\", y=\"Macro Average\", title=\"Label Model Performance by Condition and Set\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(10, 5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:09.723423Z","iopub.execute_input":"2026-09-17T20:15:09.72377Z","iopub.status.idle":"2026-09-17T20:15:10.452137Z","shell.execute_reply.started":"2026-09-17T20:15:09.723727Z","shell.execute_reply":"2026-09-17T20:15:10.451302Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(train_test_scores_df, aes(x = 'condition', y = 'precision', fill = 'set'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = CONDITION_ORDER, labels = CONDITION_LABELS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = TRAIN_TEST_SET_COLORS,\n            limits = TRAIN_TEST_SET_ORDER,\n            name = \"Set\",\n            labels = ['Train', 'Test']\n        )\n        + labs(x=\"Condition\", y=\"Precision\", title=\"Positive Predictive Value (Precision) by Condition, Set, and Class\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:10.453491Z","iopub.execute_input":"2026-09-17T20:15:10.453857Z","iopub.status.idle":"2026-09-17T20:15:11.672713Z","shell.execute_reply.started":"2026-09-17T20:15:10.453824Z","shell.execute_reply":"2026-09-17T20:15:11.671746Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(train_test_scores_df, aes(x = 'condition', y = 'npv', fill = 'set'))\n        + geom_col(position = 'dodge', color = 'white')\n        + facet_wrap('Finding')\n        + scale_x_discrete(limits = CONDITION_ORDER, labels = CONDITION_LABELS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = TRAIN_TEST_SET_COLORS,\n            limits = TRAIN_TEST_SET_ORDER,\n            name = \"Set\",\n            labels = ['Train', 'Test']\n        )\n        + labs(x=\"Condition\", y=\"NPV\", title=\"Negative Predictive Value (NPV) by Condition, Set, and Class\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 7.5))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:11.673941Z","iopub.execute_input":"2026-09-17T20:15:11.67435Z","iopub.status.idle":"2026-09-17T20:15:14.008199Z","shell.execute_reply.started":"2026-09-17T20:15:11.674307Z","shell.execute_reply":"2026-09-17T20:15:14.007081Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#ROC Curves and AUC by Condition and Set\n\nROC_GRID = np.linspace(0, 1, 101)\n\n\ndef get_roc_curves_by_class(gold_df: pd.DataFrame, probs_df: pd.DataFrame, findings: list = FINDING_COLUMNS) -> tuple:\n    scored_df = pd.merge(\n        gold_df[['StudyInstanceUID'] + findings],\n        probs_df[['StudyInstanceUID'] + findings],\n        on = 'StudyInstanceUID',\n        suffixes = ('_gold', '_prob')\n    )\n    curve_parts, auc_parts = [], []\n    for col in findings:\n        gold, probs = scored_df[f'{col}_gold'].astype(float), scored_df[f'{col}_prob'].astype(float)\n        fpr, tpr, _ = roc_curve(gold, probs)\n        curve_parts.append(pd.DataFrame({'Finding': col, 'fpr': fpr, 'tpr': tpr}))\n        auc_parts.append({'Finding': col, 'roc_auc': roc_auc_score(gold, probs)})\n\n    # The macro average interpolates each class onto a shared false positive grid; its AUC is the\n    # mean of the class AUCs rather than the area under the interpolated curve\n    macro_curve_df = pd.DataFrame({\n        'Finding': MACRO_AVERAGE_LABEL,\n        'fpr': ROC_GRID,\n        'tpr': np.mean([np.interp(ROC_GRID, part['fpr'], part['tpr']) for part in curve_parts], axis = 0)\n    })\n    macro_auc = np.mean([part['roc_auc'] for part in auc_parts])\n\n    return (\n        pd.concat(curve_parts + [macro_curve_df], ignore_index = True),\n        pd.concat([pd.DataFrame(auc_parts), pd.DataFrame([{'Finding': MACRO_AVERAGE_LABEL, 'roc_auc': macro_auc}])], ignore_index = True)\n    )\n\n\nroc_curve_parts, roc_auc_parts = [], []\nfor condition in CONDITION_ORDER:\n    for set_name in TRAIN_TEST_SET_ORDER:\n        set_curves_df, set_auc_df = get_roc_curves_by_class(\n            train_test_model_gold[(train_test_model_gold['condition'] == condition) & (train_test_model_gold['set'] == set_name)],\n            train_test_model_probs[(train_test_model_probs['condition'] == condition) & (train_test_model_probs['set'] == set_name)]\n        )\n        roc_curve_parts.append(set_curves_df.assign(condition = condition, set = set_name))\n        roc_auc_parts.append(set_auc_df.assign(condition = condition, set = set_name))\n\nFINDING_PANEL_ORDER = FINDING_COLUMNS + [MACRO_AVERAGE_LABEL]\n\ntrain_test_roc_df = pd.concat(roc_curve_parts, ignore_index = True)\ntrain_test_roc_df['Finding'] = pd.Categorical(train_test_roc_df['Finding'], categories = FINDING_PANEL_ORDER, ordered = True)\ntrain_test_roc_df['set'] = pd.Categorical(train_test_roc_df['set'], categories = TRAIN_TEST_SET_ORDER, ordered = True)\n\ntrain_test_auc_df = pd.concat(roc_auc_parts, ignore_index = True)\ntrain_test_auc_df['Finding'] = pd.Categorical(train_test_auc_df['Finding'], categories = FINDING_PANEL_ORDER, ordered = True)\ntrain_test_auc_df['set'] = pd.Categorical(train_test_auc_df['set'], categories = TRAIN_TEST_SET_ORDER, ordered = True)\ntrain_test_auc_df.pivot(index = 'Finding', columns = ['condition', 'set'], values = 'roc_auc')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:14.009489Z","iopub.execute_input":"2026-09-17T20:15:14.009803Z","iopub.status.idle":"2026-09-17T20:15:14.263665Z","shell.execute_reply.started":"2026-09-17T20:15:14.009763Z","shell.execute_reply":"2026-09-17T20:15:14.262688Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"TRAIN_TEST_CONDITION_COLORS = {\n    'criteria': '#1f77b4',\n    'consensus': '#7b3294',\n}\n\n(\n    ggplot(train_test_roc_df, aes(x = 'fpr', y = 'tpr', color = 'condition', linetype = 'set'))\n        + geom_abline(intercept = 0, slope = 1, color = '#878f99', linetype = 'dashed', size = 0.4)\n        + geom_step(direction = 'hv')\n        + facet_wrap('Finding', ncol = 4)\n        + scale_x_continuous(limits = [0,1], breaks = [0, 0.5, 1])\n        + scale_y_continuous(limits = [0,1], breaks = [0, 0.5, 1])\n        + scale_color_manual(\n            values = TRAIN_TEST_CONDITION_COLORS,\n            limits = CONDITION_ORDER,\n            name = \"Condition\",\n            labels = CONDITION_LABELS\n        )\n        + scale_linetype_manual(\n            values = ['solid', 'dotted'],\n            limits = TRAIN_TEST_SET_ORDER,\n            name = \"Set\",\n            labels = ['Train', 'Test']\n        )\n        + labs(x=\"False Positive Rate\", y=\"True Positive Rate\", title=\"ROC Curves by Condition, Set, and Class\")\n        + theme(figure_size=(12.5, 10))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:14.264901Z","iopub.execute_input":"2026-09-17T20:15:14.265219Z","iopub.status.idle":"2026-09-17T20:15:16.168454Z","shell.execute_reply.started":"2026-09-17T20:15:14.265188Z","shell.execute_reply":"2026-09-17T20:15:16.16738Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"(\n    ggplot(train_test_auc_df, aes(x = 'condition', y = 'roc_auc', fill = 'set'))\n        + geom_col(position = 'dodge', color = 'white')\n        + geom_hline(yintercept = 0.5, color = '#878f99', linetype = 'dashed')\n        + facet_wrap('Finding', ncol = 4)\n        + scale_x_discrete(limits = CONDITION_ORDER, labels = CONDITION_LABELS)\n        + scale_y_continuous(limits = [0,1])\n        + scale_fill_manual(\n            values = TRAIN_TEST_SET_COLORS,\n            limits = TRAIN_TEST_SET_ORDER,\n            name = \"Set\",\n            labels = ['Train', 'Test']\n        )\n        + labs(x=\"Condition\", y=\"ROC AUC\", title=\"ROC AUC by Condition, Set, and Class\")\n        + theme(axis_text_x=element_text(angle=45, hjust=1),\n            figure_size=(12.5, 8))\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2026-09-17T20:15:16.169892Z","iopub.execute_input":"2026-09-17T20:15:16.170315Z","iopub.status.idle":"2026-09-17T20:15:17.635394Z","shell.execute_reply.started":"2026-09-17T20:15:16.170267Z","shell.execute_reply":"2026-09-17T20:15:17.634431Z"}},"outputs":[],"execution_count":null}]}