{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Local Search predictions based in Training labels set Shannon Entropy mimic baseline","metadata":{}},{"cell_type":"markdown","source":"The first thing to understand this baseline is: we don´t work with any $X$ (set of variables to explote, in this case this data is included in the medical images), instead we uniquely use $Y$ the labels of training set (train.csv).\n\nWe will include the metric proporcionated by RSNA, weighted log loss function. Then we use it and a basic local search algorithm implementation to search a pseudo-optimal constant for every label to predict in the new data. \n\nIn difference with other baselines, that we need to weight our predictions in order to improve the results, now it is a task solved by the metaheuristic algorithm. This may result strongly in overfit. We assume that obtein a good score with the training labels must work to test too. We assume that the \"diversity distribution\" (all Shannon entropies) in training set is similar enough to the \"diversity distribution\" (all Shannon entropies) in test set. Once this premise is wrong, this baseline nosedive.\n\nTo be formal speakers, we define the set Shannon entropy with the known:\n$$ H(X) = \\sum_{i} p(x_{i}) \\log \\left( \\frac{1}{p(x_{i})} \\right) $$\n\nIn detail, we are defining a state $k$ as a possible combination of labels from all of them. Meaning that, that \"similar enough\" means that $H(X_{in})$ the set train shannon entropy and $H(X_{test})$ the set test shannon entropy contain nearly enough the same $H$ entropies of k-states. Informally speaking, there are enough pacients in the test set with the same healthy condition already seen in the train set that this optimization also works well.\n","metadata":{}},{"cell_type":"markdown","source":"** Note: Some methods used in this notebook are reciclated from this another notebook: https://www.kaggle.com/code/jakebrusca/rsna23-weighted-mean-baseline","metadata":{}},{"cell_type":"markdown","source":"# Dependecies","metadata":{}},{"cell_type":"code","source":"import random\nimport numpy as np\nimport pandas as pd\nimport pandas.api.types\nimport sklearn.metrics\n\n# Fijar una semilla\nrandom.seed(1)","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:06:21.50522Z","iopub.execute_input":"2023-08-29T12:06:21.506042Z","iopub.status.idle":"2023-08-29T12:06:21.511717Z","shell.execute_reply.started":"2023-08-29T12:06:21.506003Z","shell.execute_reply":"2023-08-29T12:06:21.510542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading targets data","metadata":{}},{"cell_type":"code","source":"labels = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv')\nlabels.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:06:21.535286Z","iopub.execute_input":"2023-08-29T12:06:21.535707Z","iopub.status.idle":"2023-08-29T12:06:21.560659Z","shell.execute_reply.started":"2023-08-29T12:06:21.535674Z","shell.execute_reply":"2023-08-29T12:06:21.559552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# List of Targets\nInjuries = ['bowel_healthy', 'bowel_injury', \n            'extravasation_healthy', 'extravasation_injury', \n            'kidney_healthy', 'kidney_low', 'kidney_high', \n            'liver_healthy', 'liver_low', 'liver_high', \n            'spleen_healthy', 'spleen_low', 'spleen_high', \n            'any_injury']\n","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:06:21.565795Z","iopub.execute_input":"2023-08-29T12:06:21.566296Z","iopub.status.idle":"2023-08-29T12:06:21.570873Z","shell.execute_reply.started":"2023-08-29T12:06:21.566266Z","shell.execute_reply":"2023-08-29T12:06:21.569957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# The metric","metadata":{}},{"cell_type":"code","source":"class ParticipantVisibleError(Exception):\n    pass\n\ndef normalize_probabilities_to_one(df: pd.DataFrame, group_columns: list) -> pd.DataFrame:\n    # Normalize the sum of each row's probabilities to 100%.\n    # 0.75, 0.75 => 0.5, 0.5\n    # 0.1, 0.1 => 0.5, 0.5\n    row_totals = df[group_columns].sum(axis=1)\n    if row_totals.min() == 0:\n        raise ParticipantVisibleError('All rows must contain at least one non-zero prediction')\n    for col in group_columns:\n        df[col] /= row_totals\n    return df\n\n\ndef score(solution: pd.DataFrame, submission: pd.DataFrame, row_id_column_name: str) -> float:\n    '''\n    Pseudocode:\n    1. For every label group (liver, bowel, etc):\n        - Normalize the sum of each row's probabilities to 100%.\n        - Calculate the sample weighted log loss.\n    2. Derive a new any_injury label by taking the max of 1 - p(healthy) for each label group\n    3. Calculate the sample weighted log loss for the new label group\n    4. Return the average of all of the label group log losses as the final score.\n    '''\n    del solution[row_id_column_name]\n    del submission[row_id_column_name]\n\n    # Run basic QC checks on the inputs\n    if not pandas.api.types.is_numeric_dtype(submission.values):\n        raise ParticipantVisibleError('All submission values must be numeric')\n\n    if not np.isfinite(submission.values).all():\n        raise ParticipantVisibleError('All submission values must be finite')\n\n    if solution.min().min() < 0:\n        raise ParticipantVisibleError('All labels must be at least zero')\n    if submission.min().min() < 0:\n        raise ParticipantVisibleError('All predictions must be at least zero')\n\n    # Calculate the label group log losses\n    binary_targets = ['bowel', 'extravasation']\n    triple_level_targets = ['kidney', 'liver', 'spleen']\n    all_target_categories = binary_targets + triple_level_targets\n\n    label_group_losses = []\n    for category in all_target_categories:\n        if category in binary_targets:\n            col_group = [f'{category}_healthy', f'{category}_injury']\n        else:\n            col_group = [f'{category}_healthy', f'{category}_low', f'{category}_high']\n\n        solution = normalize_probabilities_to_one(solution, col_group)\n\n        for col in col_group:\n            if col not in submission.columns:\n                raise ParticipantVisibleError(f'Missing submission column {col}')\n        submission = normalize_probabilities_to_one(submission, col_group)\n        label_group_losses.append(\n            sklearn.metrics.log_loss(\n                y_true=solution[col_group].values,\n                y_pred=submission[col_group].values,\n                sample_weight=solution[f'{category}_weight'].values\n            )\n        )\n\n    # Derive a new any_injury label by taking the max of 1 - p(healthy) for each label group\n    healthy_cols = [x + '_healthy' for x in all_target_categories]\n    any_injury_labels = (1 - solution[healthy_cols]).max(axis=1)\n    any_injury_predictions = (1 - submission[healthy_cols]).max(axis=1)\n    any_injury_loss = sklearn.metrics.log_loss(\n        y_true=any_injury_labels.values,\n        y_pred=any_injury_predictions.values,\n        sample_weight=solution['any_injury_weight'].values\n    )\n\n    label_group_losses.append(any_injury_loss)\n    return np.mean(label_group_losses)","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:06:21.594153Z","iopub.execute_input":"2023-08-29T12:06:21.595037Z","iopub.status.idle":"2023-08-29T12:06:21.611798Z","shell.execute_reply.started":"2023-08-29T12:06:21.595Z","shell.execute_reply":"2023-08-29T12:06:21.610657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating the real solution adding its weights ","metadata":{}},{"cell_type":"markdown","source":"In order to evaluate the score, the appropriate weights need to be assigned. The score function defined above expects sample weights for each category of injury (bowel, extravasation, kidney, liver, spleen, any). The sample weight is assigned based on the true target for a given category. For example, if a sample has a low grade kidney injury, then \n\n    [kidney_healthy, kidney_low, kidney_high] = [0,1,0]\n\nwe would set kidney_weight = 2 for that sample. We create a dataset with some addicional columns that represets the weights.","metadata":{}},{"cell_type":"code","source":"# Assign the appropriate weights to each category\ndef create_training_solution(y_train):\n    sol_train = y_train.copy()\n    \n    # bowel healthy|injury sample weight = 1|2\n    sol_train['bowel_weight'] = np.where(sol_train['bowel_injury'] == 1, 2, 1)\n    \n    # extravasation healthy/injury sample weight = 1|6\n    sol_train['extravasation_weight'] = np.where(sol_train['extravasation_injury'] == 1, 6, 1)\n    \n    # kidney healthy|low|high sample weight = 1|2|4\n    sol_train['kidney_weight'] = np.where(sol_train['kidney_low'] == 1, 2, np.where(sol_train['kidney_high'] == 1, 4, 1))\n    \n    # liver healthy|low|high sample weight = 1|2|4\n    sol_train['liver_weight'] = np.where(sol_train['liver_low'] == 1, 2, np.where(sol_train['liver_high'] == 1, 4, 1))\n    \n    # spleen healthy|low|high sample weight = 1|2|4\n    sol_train['spleen_weight'] = np.where(sol_train['spleen_low'] == 1, 2, np.where(sol_train['spleen_high'] == 1, 4, 1))\n    \n    # any healthy|injury sample weight = 1|6\n    sol_train['any_injury_weight'] = np.where(sol_train['any_injury'] == 1, 6, 1)\n    return sol_train","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:06:21.629143Z","iopub.execute_input":"2023-08-29T12:06:21.629552Z","iopub.status.idle":"2023-08-29T12:06:21.639307Z","shell.execute_reply.started":"2023-08-29T12:06:21.629497Z","shell.execute_reply":"2023-08-29T12:06:21.638104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Making the predictions dataframe","metadata":{}},{"cell_type":"code","source":"# Functions to create neighbors and new neighborhoods\n\ndef neighborhood(vecindario):\n    distribution = np.random.normal(0.0, np.sqrt(2.0), len(vecindario))\n    nuevo_vecindario = np.clip(np.round(vecindario + distribution, 0), 1, 30)\n    return nuevo_vecindario\n\ndef neighbor(vecino):\n    distribution = lambda: random.normalvariate(0.0, np.sqrt(2.0))\n    componente_mutar = random.randint(0, len(vecino) - 1)\n    \n    nuevo_vecino = vecino.copy()\n    nuevo_vecino[componente_mutar] = np.clip(np.round(vecino[componente_mutar] + distribution(), 0), 1, 30)\n    \n    return nuevo_vecino","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:06:21.657158Z","iopub.execute_input":"2023-08-29T12:06:21.65754Z","iopub.status.idle":"2023-08-29T12:06:21.66523Z","shell.execute_reply.started":"2023-08-29T12:06:21.657497Z","shell.execute_reply":"2023-08-29T12:06:21.664082Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_evaluations = 14000 # We allow like  proves to every variable\n\ndef local_search():\n    # Constant dataframe solution with weights\n    solution = create_training_solution(labels)\n    \n    # Mean of labels\n    pred = labels.copy()\n    pred[Injuries] = labels[Injuries].mean().tolist()\n    \n    # Declaration loop variables\n    best_score = float('inf')\n        \n    for i in range(0, 100): # We do multi-starts 50 times!\n        print( 'In inicialization number: ', i)\n        \n        pesos = [random.uniform(1, 30) for _ in range(len(Injuries))]\n        \n        #Declaration loop variables\n        mejora = False\n        evaluations = 0\n        generated_neighbors = 0\n\n        while(evaluations < max_evaluations and generated_neighbors < 5*14 ): # 20 * 14 important variables\n                \n            p_solution = solution.copy()\n\n            if(mejora or generated_neighbors > 14):\n                pesos = neighborhood(pesos)\n                mejora = False\n            else:\n                pesos = neighbor(pesos)\n                \n            pred_final = pred.copy()\n            pred_final[Injuries] *= pesos\n            actual_score = score(p_solution, pred_final, 'patient_id')\n\n            if(actual_score < best_score):\n                #Change pesos and score\n                best_score = actual_score\n                best_pred = pred_final.copy()\n                print('Best current score: ', best_score)\n\n                #Refresh variables\n                generated_neighbors = 0\n                mejora = True\n\n            evaluations += 1\n            generated_neighbors += 1\n    \n    return best_pred\n    ","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:06:21.684335Z","iopub.execute_input":"2023-08-29T12:06:21.685061Z","iopub.status.idle":"2023-08-29T12:06:21.696584Z","shell.execute_reply.started":"2023-08-29T12:06:21.685017Z","shell.execute_reply":"2023-08-29T12:06:21.695589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Experiment","metadata":{}},{"cell_type":"code","source":"# Real solution and weights\nsolution_train = create_training_solution(labels)\n\n# predict a constant using Genetic search optimization \nfinal_pred = labels.copy()\nfinal_pred[Injuries] = local_search()\n\n# Score\nno_scale_score = score(solution_train,final_pred,'patient_id')\nprint(f'Training score: {no_scale_score}')","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:06:21.710937Z","iopub.execute_input":"2023-08-29T12:06:21.711321Z","iopub.status.idle":"2023-08-29T12:10:08.486Z","shell.execute_reply.started":"2023-08-29T12:06:21.71129Z","shell.execute_reply":"2023-08-29T12:10:08.484371Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"# Load submission template \nsubmission = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/sample_submission.csv')\n\n# Set output to mean of training data\nsubmission[Injuries] = final_pred[Injuries]\n\n# Save Submission!\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-08-29T12:10:08.487316Z","iopub.status.idle":"2023-08-29T12:10:08.488492Z","shell.execute_reply.started":"2023-08-29T12:10:08.4882Z","shell.execute_reply":"2023-08-29T12:10:08.488229Z"},"trusted":true},"execution_count":null,"outputs":[]}]}