{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\n#import numpy as np # linear algebra\n#import pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\n#import os\n#for dirname, _, filenames in os.walk('/kaggle/input'):\n #   for filename in filenames:\n  #      print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-02-12T15:52:51.905468Z","iopub.execute_input":"2023-02-12T15:52:51.905886Z","iopub.status.idle":"2023-02-12T15:52:51.911478Z","shell.execute_reply.started":"2023-02-12T15:52:51.905851Z","shell.execute_reply":"2023-02-12T15:52:51.910328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Date preparation and exploration**","metadata":{}},{"cell_type":"markdown","source":"let's begin by importing the neccassry libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:52:51.914071Z","iopub.execute_input":"2023-02-12T15:52:51.914779Z","iopub.status.idle":"2023-02-12T15:52:51.926131Z","shell.execute_reply.started":"2023-02-12T15:52:51.914703Z","shell.execute_reply":"2023-02-12T15:52:51.925055Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"loading dataset","metadata":{}},{"cell_type":"code","source":"RSNA_data_path = '/kaggle/input/rsna-breast-cancer-detection/train_images'\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:52:51.928368Z","iopub.execute_input":"2023-02-12T15:52:51.92881Z","iopub.status.idle":"2023-02-12T15:52:51.937136Z","shell.execute_reply.started":"2023-02-12T15:52:51.928768Z","shell.execute_reply":"2023-02-12T15:52:51.935913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#loading the training data\ntrain_df = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:52:51.938829Z","iopub.execute_input":"2023-02-12T15:52:51.939309Z","iopub.status.idle":"2023-02-12T15:52:52.014001Z","shell.execute_reply.started":"2023-02-12T15:52:51.939277Z","shell.execute_reply":"2023-02-12T15:52:52.012693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **EDA**\nLet's get the general understanding of the data","metadata":{}},{"cell_type":"code","source":"print(\"Get the first 5 rows of the data\")\nprint(\"\")\n\n# Get the first 5 rows of the data\nprint(train_df.head())\n\nprint(\"\")\nprint(\"\")\n\n# Get information about the dataframe, including data types and missing values\nprint(\"Get information about the dataframe, including data types and missing values\")\nprint(\"\")\nprint(train_df.info())\n\nprint(\"\")\nprint(\"\")\n\n# Get summary statistics for the numerical columns\nprint(\"Get summary statistics for the numerical columns\")\nprint(\"\")\n\nprint(train_df.describe())\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:52:52.016391Z","iopub.execute_input":"2023-02-12T15:52:52.016772Z","iopub.status.idle":"2023-02-12T15:52:52.090653Z","shell.execute_reply.started":"2023-02-12T15:52:52.016736Z","shell.execute_reply":"2023-02-12T15:52:52.089682Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We were able to view the first five rows of the data and get information about the dataframe including data types and missing values.\n\nWe also generated summary statistics for the numerical columns, which gives us an idea of the distribution of the data.\n\nFrom the summary statistics, we can see that the average age of the patients is 58 years with a standard deviation of 10 years, and the average BIRADS score is 0.77.\n\nNow that we have a general understanding of the data, we can move on to the next step of EDA, which is to analyze the distribution and relationships of the data","metadata":{}},{"cell_type":"markdown","source":"**Let's visualize the distribution of each feature:**","metadata":{}},{"cell_type":"code","source":"\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Plot histograms of each numerical column\ntrain_df.hist(bins=50, figsize=(20,15))\nplt.show()\n\n# Plot box plots of each numerical column\ntrain_df.plot(kind='box', subplots=True, layout=(2, 5), sharex=False, sharey=False, figsize=(20,15))\nplt.show()\n\n# Plot scatter plots of each pair of numerical columns\nsns.pairplot(train_df)\nplt.show()\n\n# Compute the correlation matrix\ncorr = train_df.corr()\nsns.heatmap(corr, cmap='coolwarm', annot=True)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:52:52.091755Z","iopub.execute_input":"2023-02-12T15:52:52.092725Z","iopub.status.idle":"2023-02-12T15:53:11.350441Z","shell.execute_reply.started":"2023-02-12T15:52:52.092689Z","shell.execute_reply":"2023-02-12T15:53:11.346753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**The above error message is indicating a problem with the seaborn library. More specifically, the error is occurring in the _unsigned_subtract function within the numpy library.\n\nThe KeyError is raised because the type of an input array (numpy.bool_) is not found in the signed_to_unsigned dictionary. signed_to_unsigned is used to map signed data types to unsigned data types, which are typically used for binary operations.\n\nIn this case, it seems that the training data contains a boolean type column, which is not supported by the _unsigned_subtract function. To resolve this issue, we will have to remove the boolean type column from the input data before passing it to the sns.pairplot function. \nThe columns \"cancer\", \"biopsy\", and \"invasive\" contain boolean values, indicating whether or not the breast was positive for malignant cancer, whether or not a follow-up biopsy was performed on the breast, and if the breast is positive for cancer, whether or not the cancer proved to be invasive, respectively which make them the core of this project**","metadata":{}},{"cell_type":"markdown","source":"**So let's use correlation Matrix and visualize using heatmaps t**","metadata":{}},{"cell_type":"code","source":"\n# Calculate the correlation matrix\ncorr = train_df.corr()\n\n# Print the correlation matrix\nprint(corr)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:55:06.908564Z","iopub.execute_input":"2023-02-12T15:55:06.908998Z","iopub.status.idle":"2023-02-12T15:55:06.944184Z","shell.execute_reply.started":"2023-02-12T15:55:06.908965Z","shell.execute_reply":"2023-02-12T15:55:06.9432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**EXPLANATION**\n\nThis appears to be a correlation matrix, which shows the relationship between various variables. Each value in the matrix represents the correlation between two variables, with values ranging from -1 (perfect negative correlation) to 1 (perfect positive correlation). Some key points to summarize:\n\nThe \"site_id\" has a strong negative correlation (-0.446908) with \"machine_id\" and a moderate negative correlation (-0.430987) with \"BIRADS\".\n\nThe \"patient_id\" has a low positive correlation (0.001124) with \"image_id\", and low negative correlations with other variables.\nThe \"image_id\" has low positive correlations with \"age\" (0.000223) and \"cancer\" (0.000223), and low negative correlations with other variables.\n\nThe \"age\" has a moderate positive correlation (0.075155) with \"cancer\" and moderate negative correlations with other variables.\n\nThe \"biopsy\" has a strong positive correlation (0.613872) with \"cancer\" and moderate positive correlations with \"invasive\" (0.514311) and \"difficult_negative_case\" (0.323064).\n\nThe \"invasive\" has a strong positive correlation (0.837815) with \"cancer\" and moderate negative correlations with \"BIRADS\" (-0.172750) and \"difficult_negative_case\" (-0.049884).\n\nThe \"BIRADS\" has a moderate positive correlation (0.181352) with \"machine_id\" and a strong negative correlation (-0.833624) with \"difficult_negative_case\".\n\nThe \"implant\" has a low positive correlation (0.018106) with \"machine_id\" and low positive correlation (0.021065) with \"difficult_negative_case\"","metadata":{}},{"cell_type":"markdown","source":"**BELOW IS A HEATMAP FOR THE ABOVE CORRELATION MATRIX**","metadata":{}},{"cell_type":"code","source":"# Create the heatmap\nplt.figure(figsize=(10,10))\nsns.heatmap(df_train.corr(), annot=True, cmap='coolwarm', fmt='.2f', linewidths=.5, annot_kws={'size': 12})\nplt.title('Correlation Matrix', fontsize=15)\nplt.show()\n\n\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:55:11.98531Z","iopub.execute_input":"2023-02-12T15:55:11.985726Z","iopub.status.idle":"2023-02-12T15:55:12.705644Z","shell.execute_reply.started":"2023-02-12T15:55:11.98569Z","shell.execute_reply":"2023-02-12T15:55:12.70452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Now let's investigate relationships between some of the features\n\nTo investigate the relationships between features in the dataset, we can use pairplot from the seaborn library. Pairplot creates a scatterplot matrix of all the columns in a dataframe, showing the relationship between each pair of columns. This can help us identify any correlations or patterns in the data.\n\nHowever, before we create the pairplot, we need to select the features that we want to show the relationships between. Based on the description of the dataset, some of the features that could potentially have a relationship with one another include:\n\nage and laterality\nage and view\nage and cancer\nimplant and laterality\nimplant and cancer\ndensity and cancer\nBIRADS and cancer\nbiopsy and invasive","metadata":{}},{"cell_type":"markdown","source":"**1. AGE AND Laterality**","metadata":{}},{"cell_type":"code","source":"# Create a pairplot with age and laterality as the two features\nsns.pairplot(df_train, x_vars=[\"age\"], y_vars=[\"laterality\"], hue=\"laterality\", height=5)\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:55:17.625709Z","iopub.execute_input":"2023-02-12T15:55:17.626353Z","iopub.status.idle":"2023-02-12T15:55:19.40304Z","shell.execute_reply.started":"2023-02-12T15:55:17.626304Z","shell.execute_reply":"2023-02-12T15:55:19.401677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**2.  AGE AND CANCER**","metadata":{}},{"cell_type":"code","source":"\n# Randomly sample a fraction of the records\nfrac = 0.1 # change this fraction to control the number of records plotted\ndf_sample = df_train.sample(frac=frac, random_state=1)\nsns.pairplot(df_train, x_vars=[\"age\"], y_vars=[\"cancer\"], height=9, aspect=1.0)\n\n# Create the scatter plot\nplt.scatter(df_sample['age'], df_sample['cancer'])\nplt.xlabel(\"Age\")\nplt.ylabel(\"Cancer (0=No, 1=Yes)\")\nplt.title(\"Relationship between Age and Cancer\")\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:55:40.966408Z","iopub.execute_input":"2023-02-12T15:55:40.966854Z","iopub.status.idle":"2023-02-12T15:55:41.368724Z","shell.execute_reply.started":"2023-02-12T15:55:40.966813Z","shell.execute_reply":"2023-02-12T15:55:41.367893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **CHECKING FOR CLASS IMBALANCE**\nTo check for class imbalance in this dataset, we need to determine the ratio of positive (breast with cancer) to negative (breast without cancer) cases. To do this, we need to count the number of positive cases and the number of negative cases in the \"cancer\" column of the training set (train.csv), and then calculate the ratio of positive cases to negative cases.\n\nIf the ratio is highly skewed towards one class, it can indicate class imbalance, which can have a significant impact on the performance of a machine learning model. To mitigate class imbalance, we may need to resample the data, for example, by oversampling the minority class, or by undersampling the majority class. we can also try using techniques like cost-sensitive learning, or using a different performance metric, such as F1-score, that is less sensitive to class imbalance.","metadata":{}},{"cell_type":"markdown","source":"1. Count the number of instances of each class in the target column:\n","metadata":{}},{"cell_type":"code","source":"class_counts = df_train['cancer'].value_counts()\nprint(class_counts)","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:55:49.907869Z","iopub.execute_input":"2023-02-12T15:55:49.908236Z","iopub.status.idle":"2023-02-12T15:55:49.916756Z","shell.execute_reply.started":"2023-02-12T15:55:49.908206Z","shell.execute_reply":"2023-02-12T15:55:49.915479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*the output shows that there are two classes: 0 and 1.*\n*0: 53548 instances*\n*1: 1158 instances*\n\n*This output indicates that there is a class imbalance in the data, with the majority of the instances (53548) belonging to class 0, and a much smaller number of instances (1158) belonging to class 1. This could impact the performance of the machine learning models if not addressed, as models may tend to predict the majority class more often, leading to a higher number of false negatives (i.e., cases of cancer that are not detected) in this scenario.*","metadata":{}},{"cell_type":"markdown","source":"2. Calculate the proportion of instances for each class:\n","metadata":{}},{"cell_type":"code","source":"class_proportions = class_counts / df_train.shape[0]\nprint(class_proportions)","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:55:54.625302Z","iopub.execute_input":"2023-02-12T15:55:54.62603Z","iopub.status.idle":"2023-02-12T15:55:54.63401Z","shell.execute_reply.started":"2023-02-12T15:55:54.62599Z","shell.execute_reply":"2023-02-12T15:55:54.632664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*The output \"class_proportions = class_counts / df_train.shape[0]\nprint(class_proportions)\" displays the proportion of each class in the training data. In this case, the two classes are \"cancer\" and \"no cancer\".*\n\n*The first class, \"0\" (no cancer), is present in 97.88% of the training data.*\n*The second class, \"1\" (cancer), is present in 2.12% of the training data.*\n*This suggests that the dataset has a high class imbalance, with a large proportion of the data belonging to the \"no cancer\" class, and only a small proportion belonging to the \"cancer\" class. This class imbalance can impact the performance of machine learning models, as the model may be biased towards the majority class.\n*","metadata":{}},{"cell_type":"markdown","source":"3. Plot the class proportions using a bar chart to visualize the class imbalance","metadata":{}},{"cell_type":"code","source":"plt.bar(class_proportions.index, class_proportions)\nplt.xlabel(\"Class\")\nplt.ylabel(\"Proportion\")\nplt.title(\"Class Imbalance\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:55:58.501507Z","iopub.execute_input":"2023-02-12T15:55:58.501909Z","iopub.status.idle":"2023-02-12T15:55:58.963213Z","shell.execute_reply.started":"2023-02-12T15:55:58.501876Z","shell.execute_reply":"2023-02-12T15:55:58.962281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(class_proportions)","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:53:11.367805Z","iopub.status.idle":"2023-02-12T15:53:11.368349Z","shell.execute_reply.started":"2023-02-12T15:53:11.368066Z","shell.execute_reply":"2023-02-12T15:53:11.368092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **DATA PRE-PROCESSING**","metadata":{}},{"cell_type":"markdown","source":"1. **Data cleaning and missing value handling**","metadata":{}},{"cell_type":"markdown","source":"*checking for missing values*","metadata":{}},{"cell_type":"code","source":"# Check for missing values\nprint(train_df.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:53:11.369959Z","iopub.status.idle":"2023-02-12T15:53:11.370508Z","shell.execute_reply.started":"2023-02-12T15:53:11.370217Z","shell.execute_reply":"2023-02-12T15:53:11.370243Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From this output, it can be seen that there are missing values in the \"age\", \"BIRADS\", and \"density\" features. These missing values will need to be handled in a later step of the data preprocessing","metadata":{}},{"cell_type":"markdown","source":"**Now let\"s replace the missing values of the Age, BIRADS and Density column with their mean values**","metadata":{}},{"cell_type":"code","source":"train_df.dropna(inplace=True)\nprint(train_df.isnull().sum())","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:56:05.186926Z","iopub.execute_input":"2023-02-12T15:56:05.187297Z","iopub.status.idle":"2023-02-12T15:56:05.215911Z","shell.execute_reply.started":"2023-02-12T15:56:05.187264Z","shell.execute_reply":"2023-02-12T15:56:05.214693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Note that the inplace argument is set to True, which means that the* *original dataframe will be modified and no new dataframe will be* *created. The code above will remove all rows that have at least one* *missing value. *","metadata":{}},{"cell_type":"markdown","source":"**2. checking for incorrect values**","metadata":{}},{"cell_type":"markdown","source":"# age","metadata":{}},{"cell_type":"markdown","source":"**This code removes any rows from the train_df dataframe where the age is less than 0. If there are no acceptable values or unexpected patterns to check for in a column, the code is not necessary**","metadata":{}},{"cell_type":"code","source":"# Check for values that are outside of a specified range\ntrain_df = train_df[train_df.age >= 0]","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:56:16.047108Z","iopub.execute_input":"2023-02-12T15:56:16.047513Z","iopub.status.idle":"2023-02-12T15:56:16.056948Z","shell.execute_reply.started":"2023-02-12T15:56:16.047477Z","shell.execute_reply":"2023-02-12T15:56:16.055891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**This code first creates a list of acceptable values for the laterality column, then creates a Boolean mask by applying the isin() method to the laterality column and negating the result with the ~ operator. The value_counts() method is then used to print the counts of unique values that are not in the set of acceptable values.**","metadata":{}},{"cell_type":"markdown","source":"# laterality","metadata":{}},{"cell_type":"code","source":"# Set of acceptable values for laterality\nacceptable_laterality = {\"L\", \"R\"}\n\n# Check if there are any values in the laterality column that are not in the set of acceptable values\nincorrect_laterality = train_df[~train_df['laterality'].isin(acceptable_laterality)].index\n\n# Print the number of incorrect values found\nprint(\"Number of incorrect values in laterality:\", len(incorrect_laterality))\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:56:21.777237Z","iopub.execute_input":"2023-02-12T15:56:21.777644Z","iopub.status.idle":"2023-02-12T15:56:21.786055Z","shell.execute_reply.started":"2023-02-12T15:56:21.777595Z","shell.execute_reply":"2023-02-12T15:56:21.785221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# view","metadata":{}},{"cell_type":"code","source":"# Define a set of acceptable values for the \"view\" column\naccepted_views = {\"CC\", \"MLO\"}\n\n# Check if each value in the \"view\" column is in the set of acceptable values\nview_is_valid = train_df[\"view\"].isin(accepted_views)\n\n# Count the number of invalid values in the \"view\" column\ninvalid_view_count = sum(~view_is_valid)\n\n# Print the number of invalid values in the \"view\" column\nprint(f\"Number of invalid values in the 'view' column: {invalid_view_count}\")\n\n# Drop the rows with invalid values in the \"view\" column\ntrain_df = train_df[view_is_valid]\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:56:26.066835Z","iopub.execute_input":"2023-02-12T15:56:26.067254Z","iopub.status.idle":"2023-02-12T15:56:26.083547Z","shell.execute_reply.started":"2023-02-12T15:56:26.067221Z","shell.execute_reply":"2023-02-12T15:56:26.082252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# density","metadata":{}},{"cell_type":"code","source":"# Define the set of acceptable values\naccepted_density_values = {\"A\", \"B\", \"C\", \"D\"}\n\n# Check if each value in the \"density\" column is in the set of acceptable values\ndensity_column = train_df[\"density\"]\nfor value in density_column:\n    if value not in accepted_density_values:\n        print(f\"Unexpected value in density column: {value}\")\n\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:56:30.368113Z","iopub.execute_input":"2023-02-12T15:56:30.36852Z","iopub.status.idle":"2023-02-12T15:56:30.381438Z","shell.execute_reply.started":"2023-02-12T15:56:30.368488Z","shell.execute_reply":"2023-02-12T15:56:30.380114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"3. **checking for unexpected patterns**","metadata":{}},{"cell_type":"code","source":"def plot_box_plot(column):\n    plt.figure()\n    plt.boxplot(train_df[column])\n    plt.title(\"Box Plot of {}\".format(column))\n    plt.show()\n    \nplot_box_plot(\"age\")\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:56:34.127474Z","iopub.execute_input":"2023-02-12T15:56:34.128708Z","iopub.status.idle":"2023-02-12T15:56:34.310508Z","shell.execute_reply.started":"2023-02-12T15:56:34.128652Z","shell.execute_reply":"2023-02-12T15:56:34.309213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check for unexpected patterns in BIRADS column\nif ((train_df['BIRADS'] < 0).any() or (train_df['BIRADS'] > 5).any()):\n    print(\"Unexpected values found in BIRADS column\")\nelse:\n    print(\"No unexpected values found in BIRADS column\")\n\n# Check for unexpected patterns in machine_id column\nif ((train_df['machine_id'] < 1).any()):\n    print(\"Unexpected values found in machine_id column\")\nelse:\n    print(\"No unexpected values found in machine_id column\")\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:56:43.600668Z","iopub.execute_input":"2023-02-12T15:56:43.601085Z","iopub.status.idle":"2023-02-12T15:56:43.611073Z","shell.execute_reply.started":"2023-02-12T15:56:43.601049Z","shell.execute_reply":"2023-02-12T15:56:43.609617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"check for unexpected patterns in the \"age\", \"BIRADS\", and \"machine_id\" columns, we can use a combination of visualization and statistical tests. Here's how you can do it:\n\nNote that the Z-score threshold of 3 is a common value used to identify outliers, but it can be adjusted based on the specific needs of your data.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport numpy as np\n\n# age column\nplt.hist(train_df['age'], bins=30)\nplt.show()\n\nmean = np.mean(train_df['age'])\nstd = np.std(train_df['age'])\n\nz = (train_df['age'] - mean) / std\noutliers = train_df[np.abs(z) > 3]\nprint(f\"Outliers in age column: {outliers}\")\n\n# BIRADS column\nplt.hist(train_df['BIRADS'], bins=30)\nplt.show()\n\nmean = np.mean(train_df['BIRADS'])\nstd = np.std(train_df['BIRADS'])\n\nz = (train_df['BIRADS'] - mean) / std\noutliers = train_df[np.abs(z) > 3]\nprint(f\"Outliers in BIRADS column: {outliers}\")\n\n# machine_id column\nplt.hist(train_df['machine_id'], bins=30)\nplt.show()\n\nmean = np.mean(train_df['machine_id'])\nstd = np.std(train_df['machine_id'])\n\nz = (train_df['machine_id'] - mean) / std\noutliers = train_df[np.abs(z) > 3]\nprint(f\"Outliers in machine_id column: {outliers}\")\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:56:58.828803Z","iopub.execute_input":"2023-02-12T15:56:58.829218Z","iopub.status.idle":"2023-02-12T15:56:59.51267Z","shell.execute_reply.started":"2023-02-12T15:56:58.829176Z","shell.execute_reply":"2023-02-12T15:56:59.511516Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is a description of the outliers in the dataset, which includes several columns such as \"age\", \"BIRADS\", and \"machine_id\". The data was divided into three sections, each of which describes the outliers in one of the columns: \"age\", \"BIRADS\", and \"machine_id\".\n\nFor the \"age\" column, there are 98 rows of data that have an age of 89.0. For the \"BIRADS\" column, there are 2261 rows of data with a BIRADS score of 2.0. And for the \"machine_id\" column, there are 17 rows of data with the machine id 49.\n\nIt is worth noting that all of these outliers have the same values in the ot\nher columns, such as \"cancer\" and \"invasive\", which are both 0. This information may suggest that there is some sort of data error or that these cases are duplicates.","metadata":{}},{"cell_type":"code","source":"\nfrom sklearn.preprocessing import LabelEncoder\n\n#Load the data into a pandas dataframe\n\ntrain_df = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\n\n# Initialize a label encoder\nle = LabelEncoder()\n\n# Apply the label encoder to each column with string values\ntrain_df = train_df.apply(le.fit_transform)\n\n# Save the encoded dataframe to a new csv file\ntrain_df.to_csv('encoded_data.csv', index=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:59:37.574987Z","iopub.execute_input":"2023-02-12T15:59:37.575422Z","iopub.status.idle":"2023-02-12T15:59:37.864673Z","shell.execute_reply.started":"2023-02-12T15:59:37.575376Z","shell.execute_reply":"2023-02-12T15:59:37.863764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# FEATURE SELECTION","metadata":{}},{"cell_type":"markdown","source":"we will be using the Univariate feature selection method. This method selects the top k features based on their univariate statistical scores, such as chi-squared, ANOVA F-value, mutual information, etc","metadata":{}},{"cell_type":"markdown","source":"\nHere we will use the SelectKBest class from scikit-learn's feature_selection module to perform feature selection. The f_classif function is used as the scoring function, which is appropriate for binary classification problems like this one. The k parameter is set to 10, meaning that the top 10 features will be selected.\n\n]","metadata":{}},{"cell_type":"code","source":"\nfrom sklearn.feature_selection import SelectKBest, f_classif\n\n\n\n# Define the features (predictors) and target variable\nfeatures = train_df.drop([\"cancer\"], axis=1)\ntarget = train_df[\"cancer\"]\n\n# Perform feature selection using SelectKBest with f_classif\nselector = SelectKBest(f_classif, k=10)\nselected_features = selector.fit_transform(features, target)\n\n# Get the column indices of the selected features\nselected_feature_indices = selector.get_support(indices=True)\n\n# Get the names of the selected features\nselected_feature_names = features.columns[selected_feature_indices]\n\n# Print the selected features\nprint(\"Selected features:\")\nprint(selected_feature_names)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-12T15:59:42.966633Z","iopub.execute_input":"2023-02-12T15:59:42.967612Z","iopub.status.idle":"2023-02-12T15:59:42.996242Z","shell.execute_reply.started":"2023-02-12T15:59:42.967563Z","shell.execute_reply":"2023-02-12T15:59:42.995439Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}