{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## <font color='red'> RSNA Breast Cancer Detection </font> ","metadata":{}},{"cell_type":"markdown","source":"Breast cancer detection is the aim of this competition. To train the mode, screening mammograms from routine screening was used. \n\nThe quality and safety of patient care may be improved by increasing radiologists' accuracy and efficiency through the work on enhancing the automation of detection in screening mammography. Additionally, it could aid in lowering expenses and pointless medical treatments.\n","metadata":{}},{"cell_type":"markdown","source":"<font color='red'> Context </font> ","metadata":{}},{"cell_type":"markdown","source":"Breast cancer is the type of cancer that strikes women most frequently worldwide, according to the WHO. There were 2.3 million new breast cancer diagnoses and 685,000 fatalities in 2020 alone. However, since routine mammography screening was introduced by health authorities in age groups deemed at risk in the 1980s, the death rate of breast cancer in high-income nations has decreased by 40%. Your machine learning abilities may be able to speed up the process radiologists use to assess screening mammograms since early identification and treatment are essential to lowering cancer fatality rates.\n\nCurrently, screening mammography systems are expensive to operate since early diagnosis of breast cancer necessitates the expertise of highly trained human observers. This issue will probably get worse because to a radiologists shortage that is coming to various nations. A high prevalence of false positive findings is another side effect of mammography screening. This may lead to unwarranted concern, inconvenienced follow-up treatment, more imaging tests, and occasionally the requirement for tissue sample (often a needle biopsy).\n\nThe Radiological Society of North America (RSNA), which is hosting the tournament, is a nonprofit association that represents 31 radiologic subspecialties from 145 different nations. Through research, teaching, and technology innovation, RSNA supports excellence in patient care and healthcare delivery.\n\nOur participation in this contest might contribute to spreading the advantages of early detection to more people. Greater accessibility could further lower the global death rate from breast cancer.\n","metadata":{}},{"cell_type":"markdown","source":"## <font color='red'> Data Overview </font> ","metadata":{}},{"cell_type":"markdown","source":"The description of the data is shown on the following list, where is specified that some of them are exclusive for the training set and also they give an introduction to the data form:\n\n- site_id - ID code for the source hospital.\n- patient_id - ID code for the patient.\n- image_id - ID code for the image.\n- laterality - Whether the image is of the left or right breast.\n- view - The orientation of the image. The default for a screening exam is to capture two views per breast.\n- age - The patient's age in years.\n- implant - Whether or not the patient had breast implants. Site 1 only provides breast implant information at the patient level, not at the breast level.\n- density - A rating for how dense the breast tissue is, with A being the least dense and D being the most dense. Extremely dense tissue can make diagnosis more difficult. **Only provided for train**.\n- machine_id - An ID code for the imaging device.\n- cancer - Whether or not the breast was positive for malignant cancer. The target value. **Only provided for train**.\n- biopsy - Whether or not a follow-up biopsy was performed on the breast. **Only provided for train**.\n- invasive - If the breast is positive for cancer, whether or not the cancer proved to be invasive. **Only provided for train**.\n- BIRADS - 0 if the breast required follow-up, 1 if the breast was rated as negative for cancer, and 2 if the breast was rated as normal. Only provided for train.\n- prediction_id - The ID for the matching submission row. Multiple images will share the same prediction ID. **Test only**.\n- difficult_negative_case - True if the case was unusually difficult. **Only provided for train.**\n\nFrom this description, the main EDA was built as a reference for the future Cancer Detection.","metadata":{}},{"cell_type":"code","source":"#importing the libraries\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport seaborn as sn","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:54.738359Z","iopub.execute_input":"2023-02-20T21:28:54.739306Z","iopub.status.idle":"2023-02-20T21:28:55.799473Z","shell.execute_reply.started":"2023-02-20T21:28:54.739168Z","shell.execute_reply":"2023-02-20T21:28:55.797958Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#importing our cancer dataset\ndf = pd.read_csv('/kaggle/input/rsna-bcd-roi-1024x512-png-v2-dataset/train.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:55.802081Z","iopub.execute_input":"2023-02-20T21:28:55.8026Z","iopub.status.idle":"2023-02-20T21:28:56.077394Z","shell.execute_reply.started":"2023-02-20T21:28:55.802549Z","shell.execute_reply":"2023-02-20T21:28:56.076361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#find the dimensions of the data set using the panda dataset ‘shape’ attribute.\nprint(\"Cancer data set dimensions : {}\".format(df.shape))","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:56.079076Z","iopub.execute_input":"2023-02-20T21:28:56.080264Z","iopub.status.idle":"2023-02-20T21:28:56.087682Z","shell.execute_reply.started":"2023-02-20T21:28:56.08021Z","shell.execute_reply":"2023-02-20T21:28:56.086485Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.columns","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:56.09124Z","iopub.execute_input":"2023-02-20T21:28:56.091743Z","iopub.status.idle":"2023-02-20T21:28:56.102326Z","shell.execute_reply.started":"2023-02-20T21:28:56.091702Z","shell.execute_reply":"2023-02-20T21:28:56.100723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1 = df[['site_id', 'patient_id', 'image_id', 'laterality', 'view', 'age',\n       'cancer', 'biopsy', 'invasive', 'BIRADS', 'implant', 'density',\n       'machine_id', 'difficult_negative_case']]","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:56.104308Z","iopub.execute_input":"2023-02-20T21:28:56.105246Z","iopub.status.idle":"2023-02-20T21:28:56.122323Z","shell.execute_reply.started":"2023-02-20T21:28:56.105191Z","shell.execute_reply":"2023-02-20T21:28:56.12095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df1[0:5]","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:56.124409Z","iopub.execute_input":"2023-02-20T21:28:56.124962Z","iopub.status.idle":"2023-02-20T21:28:56.148698Z","shell.execute_reply.started":"2023-02-20T21:28:56.124915Z","shell.execute_reply":"2023-02-20T21:28:56.14774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Missing or Null Data points\ndf1.isnull().sum()\ndf1.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:56.150092Z","iopub.execute_input":"2023-02-20T21:28:56.151043Z","iopub.status.idle":"2023-02-20T21:28:56.18674Z","shell.execute_reply.started":"2023-02-20T21:28:56.151Z","shell.execute_reply":"2023-02-20T21:28:56.184526Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Corralattion\n# Select only numeric columns from the DataFrame\nnumeric_columns = df1.select_dtypes(include=['int64', 'float64'])\n\n# Create the correlation matrix\ncorrelation_matrix = numeric_columns.corr()\n\n# Plot the correlation matrix using seaborn\nplt.figure(figsize=(20, 15))\nsn.heatmap(correlation_matrix, annot=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:56.188933Z","iopub.execute_input":"2023-02-20T21:28:56.189982Z","iopub.status.idle":"2023-02-20T21:28:57.3621Z","shell.execute_reply.started":"2023-02-20T21:28:56.189923Z","shell.execute_reply":"2023-02-20T21:28:57.360856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Correlation is a statistical term that measures the relationship between two variables. The correlation coefficient is a number between -1 and 1, where -1 represents a perfect negative connection (when one variable grows, the other falls), 0 represents no correlation (the variables are unrelated), and 1 represents a perfect positive correlation (when one variable increases, the other increases).","metadata":{}},{"cell_type":"code","source":"# Select only numeric columns from the DataFrame\nnumeric_columns = df1.select_dtypes(include=['int64', 'float64'])\n\n# Create the correlation matrix\ncorr_matrix = numeric_columns.corr()\n\n# Sort the correlations between the \"cancer\" column and other numeric columns\ncancer_correlations = corr_matrix['cancer'].sort_values(ascending=False)\n\nprint(cancer_correlations)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:57.363393Z","iopub.execute_input":"2023-02-20T21:28:57.364016Z","iopub.status.idle":"2023-02-20T21:28:57.404058Z","shell.execute_reply.started":"2023-02-20T21:28:57.363967Z","shell.execute_reply":"2023-02-20T21:28:57.402685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It shows the correlation of cancer with another variable. The matrix demonstrates a positive relationship between cancer, invasive, and biopsy. \n","metadata":{}},{"cell_type":"markdown","source":"## <font color='red'> Data Analysis </font> ","metadata":{}},{"cell_type":"code","source":"num_patients = df1['patient_id'].nunique()\nmin_patient_age = df1['age'].min()\nmax_patient_age = df1['age'].max()\ngrouped = df1.groupby('patient_id')['cancer'].max()\nn_negative = grouped.value_counts().get(0, 0)\nn_positive = grouped.value_counts().get(1, 0)\n\nprint(f\"There are {num_patients} different patients in the train set.\\n\")\nprint(f\"The youngest patient is {int(min_patient_age)} years old.\")\nprint(f\"The oldest patient is {int(max_patient_age)} years old.\\n\")\nprint(f\"{n_negative} patients ({n_negative / num_patients:.1%}) are negative to breast cancer.\")\nprint(f\"{n_positive} patients ({n_positive / num_patients:.1%}) are positive to breast cancer.\")\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:57.409592Z","iopub.execute_input":"2023-02-20T21:28:57.410045Z","iopub.status.idle":"2023-02-20T21:28:57.433256Z","shell.execute_reply.started":"2023-02-20T21:28:57.410008Z","shell.execute_reply":"2023-02-20T21:28:57.431866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I use f-strings to format the output strings, which allows us to include variable values directly in the string. I also use the value_counts() method on the grouped object to count the number of positive and negative cases, and use the get() method to handle the case where there are no cases of one of the values. Finally, I use :.1% to format the ratios as percentages with one decimal point.","metadata":{}},{"cell_type":"code","source":"cancer_per_patient = df1.groupby(\"patient_id\")[\"cancer\"].max().values\nn_negative = (cancer_per_patient == 0).sum()\nn_positive = (cancer_per_patient == 1).sum()\n\nfig, ax = plt.subplots()\nbars = ax.bar([\"No cancer\", \"Cancer\"], [n_negative, n_positive])\nax.set(xlabel=\"\", ylabel=\"Count\", title=\"Number of patients with cancer\")\nfor bar, count in zip(bars, [n_negative, n_positive]):\n    height = bar.get_height()\n    ax.text(bar.get_x() + bar.get_width() / 2, height, count, ha=\"center\", va=\"bottom\")\nplt.show()\n\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:57.435297Z","iopub.execute_input":"2023-02-20T21:28:57.435802Z","iopub.status.idle":"2023-02-20T21:28:57.623792Z","shell.execute_reply.started":"2023-02-20T21:28:57.435754Z","shell.execute_reply":"2023-02-20T21:28:57.621719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <font color='red'> Age Distribution </font> ","metadata":{}},{"cell_type":"code","source":"ages = df1.groupby('patient_id')['age'].apply(lambda x: x.unique()[0])\ncancer_ages = df1[df1['cancer'] == 1].groupby('patient_id')['age'].apply(lambda x: x.unique()[0])\nno_cancer_ages = df1[df1['cancer'] == 0].groupby('patient_id')['age'].apply(lambda x: x.unique()[0])\n\nplt.figure(figsize=(14, 7))\n\nplt.subplot(1, 2, 1)\nsn.histplot(ages, bins=63, color='orange', kde=True)\nplt.title(\"All Patients\", fontsize=14)\nplt.xlabel(\"Age\", fontsize=12)\nplt.ylabel(\"Count\", fontsize=12)\nplt.xticks(fontsize=10)\nplt.yticks(fontsize=10)\nplt.xlim(33, 89)\n\nplt.subplot(1, 2, 2)\nsn.histplot(cancer_ages, bins=51, color='red', kde=True)\nsn.histplot(no_cancer_ages, bins=63, color='green', kde=True)\nplt.title(\"Patients with and without cancer\", fontsize=14)\nplt.xlabel(\"Age\", fontsize=12)\nplt.ylabel(\"Count\", fontsize=12)\nplt.xticks(fontsize=10)\nplt.yticks(fontsize=10)\nplt.xlim(33, 89)\nplt.legend([\"Cancer\", \"No cancer\"], fontsize=10)\n\nplt.suptitle(\"Age distribution of the patients\", fontsize=16)\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:28:57.625565Z","iopub.execute_input":"2023-02-20T21:28:57.626612Z","iopub.status.idle":"2023-02-20T21:29:00.121436Z","shell.execute_reply.started":"2023-02-20T21:28:57.626561Z","shell.execute_reply":"2023-02-20T21:29:00.120511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I build a 1x2 grid of subplots using the subplot() function and add titles, axis labels, y-axis labels, and tick labels to each subplot. The legend() function is also used to add a legend to the second subplot, which displays color-coded bars for cancer and no cancer instances. Lastly, I add a title to the whole figure using the suptitle() technique and alter the spacing between the subplots using the tight layout() method.","metadata":{}},{"cell_type":"code","source":"# Statistics\nimport pandas as pd\n\nages = df1.groupby('patient_id')['age'].apply(lambda x: x.unique()[0])\n\nstats = pd.DataFrame({\n    \"Mean\": [ages.mean()],\n    \"Std\": [ages.std()],\n    \"Q1\": [ages.quantile(0.25)],\n    \"Median\": [ages.median()],\n    \"Q3\": [ages.quantile(0.75)],\n    \"Mode\": [ages.mode()[0]]\n})\n\nformatted_stats = stats.style.format(\"{:.2f}\")\nformatted_stats.set_caption(\"Statistics for Age\")\nformatted_stats.set_properties(**{'text-align': 'center'})\n\ndisplay(formatted_stats)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:00.122598Z","iopub.execute_input":"2023-02-20T21:29:00.123585Z","iopub.status.idle":"2023-02-20T21:29:00.833053Z","shell.execute_reply.started":"2023-02-20T21:29:00.123544Z","shell.execute_reply":"2023-02-20T21:29:00.831942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sn\n\nages = df1.groupby('patient_id')['age'].apply(lambda x: x.unique()[0])\n\nplt.figure(figsize=(6, 8))\nsn.boxplot(y=ages, color='orange')\nplt.title(\"Age distribution\")\nplt.ylabel(\"Age\")\nplt.plot([0], [ages.mean()], marker='o', color='red')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:00.834608Z","iopub.execute_input":"2023-02-20T21:29:00.834997Z","iopub.status.idle":"2023-02-20T21:29:01.622392Z","shell.execute_reply.started":"2023-02-20T21:29:00.834955Z","shell.execute_reply":"2023-02-20T21:29:01.62103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Statistics for Cancer Age","metadata":{}},{"cell_type":"code","source":"# Statistics for Cancer Age\nimport pandas as pd\n\nages = df1.groupby('patient_id')['age'].apply(lambda x: x.unique()[0])\n\nstats = pd.DataFrame({\n    \"Mean\": [cancer_ages.mean()],\n    \"Std\": [cancer_ages.std()],\n    \"Q1\": [cancer_ages.quantile(0.25)],\n    \"Median\": [cancer_ages.median()],\n    \"Q3\": [cancer_ages.quantile(0.75)],\n    \"Mode\": [cancer_ages.mode()[0]],\n    \"Minimum Cancer Age\": [cancer_ages.min()]\n})\n\nformatted_stats = stats.style.format(\"{:.2f}\")\nformatted_stats.set_caption(\"Statistics for Cancer Age\")\nformatted_stats.set_properties(**{'text-align': 'center'})\n\ndisplay(formatted_stats)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:01.62416Z","iopub.execute_input":"2023-02-20T21:29:01.624543Z","iopub.status.idle":"2023-02-20T21:29:02.296098Z","shell.execute_reply.started":"2023-02-20T21:29:01.624508Z","shell.execute_reply":"2023-02-20T21:29:02.294695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sn\n\nmy_colors = [\"#f0ad4e\", \"#5cb85c\", \"#d9534f\", \"#428bca\", \"#9b59b6\", \"#34495e\"]\n\nf, (a0, a1) = plt.subplots(2, 1, gridspec_kw={'height_ratios': [3, 1]}, figsize=(14, 10))\nsn.histplot(data=cancer_ages, kde=True, color=my_colors[5], ax=a0)\n\na0.axvline(x=cancer_ages.mean(), ls=\":\", lw=2, color=\"black\")\na0.text(x=cancer_ages.mean()+1, y=0.018, s=f\"Mean: {cancer_ages.mean():.2f}\", size=17, color=\"black\", weight=\"bold\")\na0.axvline(x=cancer_ages.min(), ls=\":\", lw=2, color=\"black\")\na0.text(x=cancer_ages.min()+1, y=0.008, s=f\"Min: {cancer_ages.min()}\", size=17, color=\"black\", weight=\"bold\")\na0.axvline(x=cancer_ages.max(), ls=\":\", lw=2, color=\"black\")\na0.text(x=cancer_ages.max()-7, y=0.037, s=f\"Max: {cancer_ages.max()}\", size=17, color=\"black\", weight=\"bold\")\n\nsn.boxenplot(x=cancer_ages, ax=a1, color=my_colors[2])\na1.set(xlabel=\"Age\", ylabel=\"\")\na1.tick_params(labelsize=12)\n\nplt.suptitle(\"Cancer Positive Age Distribution\", weight=\"bold\", size=20)\na0.set(title=\"Density Plot\")\na0.set(xlabel=\"Age\", ylabel=\"Density\")\na0.tick_params(labelsize=12)\na0.spines[\"top\"].set_visible(False)\na0.spines[\"right\"].set_visible(False)\na0.spines[\"left\"].set_linewidth(2)\na0.spines[\"bottom\"].set_linewidth(2)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:02.297983Z","iopub.execute_input":"2023-02-20T21:29:02.298375Z","iopub.status.idle":"2023-02-20T21:29:02.86305Z","shell.execute_reply.started":"2023-02-20T21:29:02.29834Z","shell.execute_reply":"2023-02-20T21:29:02.861815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I use the histplot() method to generate the density plot and set the kde argument to True in order to include a kernel density estimate on the plot. I also utilize the set() function to modify the font size and weight of the subplots' content. I modify the font size of the tick labels using the tick params() function. Lastly, the tight layout() function is used to modify the distance between subplots.\n\nThe graph represents the age distribution of cancer-positive individuals. According to the age distribution, the mean age of cancer patients is older than that of people without cancer. In other words, the likelihood of developing cancer increases with age.\n\nIt is also evident from the following violin figure that cancer patients are disproportionately prevalent among the elderly (age between 60 - 70).","metadata":{}},{"cell_type":"code","source":"# Violinplot\n\ncolors = [\"lightblue\", \"lightcoral\"]\nf, ax = plt.subplots(figsize=(12, 8))\nsn.violinplot(x=\"cancer\", y=\"age\", data=df1, palette=colors, ax=ax)\n\nax.set_xlabel(\"Cancer\", fontsize=14, fontweight=\"bold\")\nax.set_ylabel(\"Age\", fontsize=14, fontweight=\"bold\")\nax.tick_params(labelsize=12)\n\nplt.title(\"Age Distribution by Cancer Status\", fontsize=20, fontweight=\"bold\")\nsn.despine(trim=True, left=True)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:02.864598Z","iopub.execute_input":"2023-02-20T21:29:02.865892Z","iopub.status.idle":"2023-02-20T21:29:03.283147Z","shell.execute_reply.started":"2023-02-20T21:29:02.865839Z","shell.execute_reply":"2023-02-20T21:29:03.281946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The palette argument is set to a list of cancer category colors, and the figsize parameter is used to alter the size of the figure. I set the labels for the x- and y-axes using the set xlabel() and set ylabel() methods, and I alter the font size of the tick labels using the tick params() method. We also set the plot title using the title() method. I then use the despine() method to remove the plot's top and right spines, and I modify the trim and left parameters to make the plot more compact.","metadata":{}},{"cell_type":"markdown","source":"# Insghts","metadata":{}},{"cell_type":"markdown","source":"We may conclude from the data that age is a significant determinant in breast cancer diagnosis. The majority of patients in the sample are older than 40 years old, with the distribution peaking at the age of 64. After the peak, the number of patients declines, which may be attributable to variables such as a decline in life expectancy or screening rates among elderly patients.\n\nCancer patients tend to be older than those without the disease. Thus, age may be a valuable characteristic for constructing a model to predict breast cancer diagnosis. Overall, the age distribution of patients gives valuable information that might aid in the development of efficient breast cancer screening and diagnostic procedures.","metadata":{}},{"cell_type":"markdown","source":"## <font color='red'> Number of Images Per patient </font> ","metadata":{}},{"cell_type":"code","source":"num_images_per_patient = df1['patient_id'].value_counts()\nplt.figure(figsize=(14, 8))\nsn.countplot(x=num_images_per_patient, palette='Reds_r', saturation=0.7)\nplt.title(\"Number of Images Taken per Patient\", fontsize=20, fontweight=\"bold\")\nplt.xlabel('Number of Images Taken', fontsize=14, fontweight=\"bold\")\nplt.ylabel('Count of Patients', fontsize=14, fontweight=\"bold\")\nplt.xticks(rotation=90)\nplt.tick_params(labelsize=12)\nplt.tight_layout()\nsn.despine(trim=True)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:03.284799Z","iopub.execute_input":"2023-02-20T21:29:03.285288Z","iopub.status.idle":"2023-02-20T21:29:03.633811Z","shell.execute_reply.started":"2023-02-20T21:29:03.285242Z","shell.execute_reply":"2023-02-20T21:29:03.632413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I alter the size of the plot by using the figsize option. In addition, I adjusted the saturation value to 0.7 to intensify the color of the bars. I set the title and labels for the plot using the title(), xlabel(), and ylabel() methods, and I rotate the x-tick labels using the rotation option. We modify the font size of the tick labels using the tick params() method. I then use the tight layout() function to alter the distance between subplots, the despine() method to eliminate the plot's top and right spines, and the trim parameter to make the plot more compact.","metadata":{}},{"cell_type":"markdown","source":"The majority of patients, according to the countplot of the number of photographs collected per patient, have four images, which corresponds to two views each side (left and right). However, there are instances in which people experience more than four pictures, which may be the result of a number of causes, including:\n\n1. Several imaging modalities: In addition to mammography, patients may have also received MRI, ultrasound, or PET scans, which may have contributed to the larger number of pictures.\n\n2. Individuals with a history of breast cancer or a high risk of acquiring the illness may be needed to undergo more regular follow-up imaging, which may result in a greater number of pictures.\n\n3. Individuals with a history of breast surgery, breast implants, or other complicating conditions may require additional scans for the accurate diagnosis of any anomalies.\n\nWhile the majority of patients had four photos, the prevalence of patients with more images underscores the difficulty of breast cancer detection and the significance of individualized imaging regimens.","metadata":{}},{"cell_type":"markdown","source":"## <font color='red'> Image Feature </font> ","metadata":{}},{"cell_type":"code","source":"fig, ax = plt.subplots(2, 3, figsize=(18, 12))\nsn.countplot(x=df1['laterality'], palette='Blues_r', ax=ax[0, 0])\nsn.countplot(x=df1['implant'], palette='Greens_r', ax=ax[0, 1])\nsn.countplot(x=df1['difficult_negative_case'], palette='Reds_r', ax=ax[0, 2])\nsn.countplot(x=df1['view'], palette='Oranges_r', ax=ax[1, 0])\nsn.countplot(x=df1['density'], palette='Purples_r', order=['A', 'B', 'C', 'D'], ax=ax[1, 1])\nsn.countplot(x=df1['site_id'], palette='Greys_r', ax=ax[1, 2])\n\nax[0, 0].set_title('Laterality', fontsize=18, fontweight='bold')\nax[0, 1].set_title('Implant', fontsize=18, fontweight='bold')\nax[0, 2].set_title('Difficult Negative Case', fontsize=18, fontweight='bold')\nax[1, 0].set_title('View', fontsize=18, fontweight='bold')\nax[1, 1].set_title('Density', fontsize=18, fontweight='bold')\nax[1, 2].set_title('Site ID', fontsize=18, fontweight='bold')\n\nfor i in range(2):\n    for j in range(3):\n        ax[i, j].set_xlabel('')\n        ax[i, j].set_ylabel('')\n        ax[i, j].tick_params(axis='both', which='major', labelsize=14)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:03.635323Z","iopub.execute_input":"2023-02-20T21:29:03.636284Z","iopub.status.idle":"2023-02-20T21:29:04.658954Z","shell.execute_reply.started":"2023-02-20T21:29:03.63624Z","shell.execute_reply":"2023-02-20T21:29:04.657997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I alter the size of the plot by using the figsize option. Also, I utilize the x argument to specify the columns of interest for each countplot. The title() method is used to set the title for each subplot, and the set xlabel() and set ylabel() procedures are used to remove the x- and y-axis labels for each subplot. We modify the font size of the tick labels using the tick params() method. Lastly, I use the tight layout() method to change the distance between subplots and the trim parameter to condense the plot.","metadata":{}},{"cell_type":"markdown","source":"On the basis of the countplots of various picture qualities, the following conclusions may be drawn:\n\n1. In terms of laterality, the photos are balanced, suggesting that the same amount of images were captured for both the left and right breast.\n2. Implants are uncommon, with just a few photos indicating their presence.\n3. Several photos were difficult to diagnose, which may indicate the existence of breast tissue that is more complicated or unusual.\n4. The bulk of photos consist of two different types of views: CC (craniocaudal) and MLO (mediolateral oblique).\n5. The density of breast tissue in the photos is often in the medium range (B and C), with a minority of photographs showing breast tissue density that is either extremely dense (D) or extremely low (A).\n6. The photos were obtained in a balanced manner at two distinct locations, which may indicate that the data came from two distinct medical institutions or imaging centers.\n\nWhile constructing a model for breast cancer diagnosis or risk assessment, it is crucial to account for various image properties and patient variables.","metadata":{}},{"cell_type":"markdown","source":"## <font color='red'> Exploring Target Features </font> ","metadata":{}},{"cell_type":"code","source":"biopsy_counts = df1.groupby(['cancer', 'biopsy']).size().unstack(fill_value=0)\nbiopsy_perc = biopsy_counts.apply(lambda x: x / x.sum(), axis=1)\n\nfig, ax = plt.subplots(1, 2, figsize=(12, 6))\nsn.countplot(x='biopsy', hue='cancer', data=df1, palette='Set2', ax=ax[0])\nax[0].bar_label(ax[0].containers[0], label_type='edge')\nax[0].bar_label(ax[0].containers[1], label_type='edge')\nsn.heatmap(biopsy_perc, square=True, annot=True, fmt='.1%', cmap='Blues', ax=ax[1], cbar=False)\nax[1].set_yticklabels(ax[1].get_yticklabels(), rotation=0)\nplt.subplots_adjust(wspace=0.3)\nax[0].set_title(\"Number of images resulting in a biopsy\")\nax[1].set_title(\"Percentage of images resulting in a biopsy\")\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:04.660147Z","iopub.execute_input":"2023-02-20T21:29:04.660902Z","iopub.status.idle":"2023-02-20T21:29:04.976546Z","shell.execute_reply.started":"2023-02-20T21:29:04.660864Z","shell.execute_reply":"2023-02-20T21:29:04.975278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The groupby function is used to both the cancer and biopsy columns, and the size function is used to determine the frequency of each combination of malignant and noncancerous photos and biopsy findings. The unstack function is used to reformat the Series into a DataFrame with one row for each combination of malignant and non-cancerous photos and one column for each biopsy result. The fill value option is set to 0 so that missing values are replaced with zeros.\n\nThe outcomeant biopsy counts The DataFrame is then utilized to compute the proportion of malignant and non-cancerous pictures that result in a biopsy. This is accomplished by dividing each entry of the biopsy counts DataFrame by its total using the apply method.\n\nExcept for the first line, which calculates biopsy counts differently than before, the remainder of the code remains same. Now, the countplot function is used to construct a bar plot showing the number of photos that resulted in a biopsy, with the x parameter set to biopsy, the hue parameter set to cancer, and the palette parameter set to 'Set2' in order to employ a color scheme that works well for two categories. Using the heatmap function as previously, a heatmap of the percentage of photographs resulting in a biopsy is generated.\n\n### Insights\nThe breast cancer dataset has a comparatively small number of sample photos with cancer compared to healthy breast samples, making it difficult for the algorithm to discover the necessary pattern. In addition, all cancer patients got a biopsy, but only a tiny fraction of non-cancerous scans result in the same operation.","metadata":{}},{"cell_type":"code","source":"#import matplotlib.pyplot as plt\n#import seaborn as sns\n\nfig, ax = plt.subplots(1, 3, figsize=(16, 4))\n\nsn.countplot(x='invasive', data=df1[df1['cancer'] == True], ax=ax[0], color='red')\nsn.countplot(x='BIRADS', data=df1[df1['cancer'] == False], order=[0, 1, 2], ax=ax[1], color='blue')\nsn.countplot(x='BIRADS', data=df1[df1['cancer'] == True], order=[0, 1, 2], ax=ax[2], color='purple')\n\nax[0].set_title(\"Count of Invasive Cancer Images\")\nax[0].set_xlabel(\"Invasive\")\nax[0].set_ylabel(\"Count\")\n\nax[1].set_title(\"BIRADS for Healthy Images\")\nax[1].set_xlabel(\"BIRADS\")\nax[1].set_ylabel(\"Count\")\n\nax[2].set_title(\"BIRADS for Cancer Images\")\nax[2].set_xlabel(\"BIRADS\")\nax[2].set_ylabel(\"Count\")\n\nplt.suptitle(\"Breast Cancer Image Analysis\", fontsize=14)\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:04.978145Z","iopub.execute_input":"2023-02-20T21:29:04.978494Z","iopub.status.idle":"2023-02-20T21:29:05.36138Z","shell.execute_reply.started":"2023-02-20T21:29:04.978462Z","shell.execute_reply":"2023-02-20T21:29:05.360236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With the x argument of the sns.countplot() function, the variable to be counted is specified. In addition, I utilize the color parameter to provide a uniform color scheme for all three subplots. The set xlabel(), set ylabel(), and set title() routines are used to add subplot labels and figure titles.\n\nBIRADS value signification:\n* 0: required follow-up\n* 1: rated as negative for cancer\n* 2: rated as normal\n\n### Insights\nThe majority of cancer photos were classified as invasive, and all cancer images warranted further investigation. Nonetheless, even some benign scans resulted in further investigation, underscoring the need for regular screenings and checks.","metadata":{}},{"cell_type":"markdown","source":"## <font color='red'> Machine ID </font> ","metadata":{}},{"cell_type":"code","source":"count_machine=df1.groupby(by=\"machine_id\").count()[\"patient_id\"]\ncount_machine.reset_index().head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:05.362879Z","iopub.execute_input":"2023-02-20T21:29:05.363317Z","iopub.status.idle":"2023-02-20T21:29:05.39396Z","shell.execute_reply.started":"2023-02-20T21:29:05.363282Z","shell.execute_reply":"2023-02-20T21:29:05.392701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,8))\nplt.title(\"Number of images per machine_id\")\nsn.barplot(data=count_machine.reset_index(),x=\"machine_id\",y=\"patient_id\", color='b')\nplt.ylabel(\"Number of images\")","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:05.395649Z","iopub.execute_input":"2023-02-20T21:29:05.39619Z","iopub.status.idle":"2023-02-20T21:29:05.672389Z","shell.execute_reply.started":"2023-02-20T21:29:05.396141Z","shell.execute_reply":"2023-02-20T21:29:05.671152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code generates a bar chart in seaborn (sn) that displays the amount of photos for each machine id. The y-axis of the graph represents the total number of photos, while the x-axis represents the machine ids.\n### Insights\nThe data suggests that 10 distinct devices were used to capture the photos, but just four machines accounted for the vast bulk of the total (49, 21, 29, and 48). It is important to account for the influence of machine variations while evaluating the picture dataset due to the fact that changes in machines might lead to variances in the distribution of images.","metadata":{}},{"cell_type":"markdown","source":"## <font color='red'> IMAGE PROCESSING </font> ","metadata":{}},{"cell_type":"markdown","source":"https://www.kaggle.com/datasets/awsaf49/rsna-bcd-roi-1024x512-png-v2-dataset \n\nThe data images were downloaded from the address. It's a PNG file with a smaller file size. \n","metadata":{}},{"cell_type":"code","source":"###Basic Libs..\nimport warnings\nwarnings.filterwarnings(\"ignore\")\nimport pandas as pd\nimport numpy as np\nfrom tqdm import tqdm,tqdm_notebook\n#from prettytable import PrettyTable\nimport pickle\nimport os\nprint('CWD is ',os.getcwd())\n\n# Vis Libs..\nfrom sklearn.manifold import TSNE\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nplt.rcParams[\"axes.grid\"] = False\n\n# Image Libs.\nfrom PIL import Image\nimport cv2\n\nimport tensorflow as tf, re, math\n\n# DL Libs..\nimport keras\nfrom keras import applications\nfrom keras.preprocessing.image import ImageDataGenerator #,img_to_array,array_to_img,load_img\nfrom keras import optimizers,Model,Sequential\nfrom keras.layers import Input,GlobalAveragePooling2D,Dropout,Dense,Activation\nfrom keras.callbacks import EarlyStopping,ReduceLROnPlateau","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:05.674287Z","iopub.execute_input":"2023-02-20T21:29:05.675041Z","iopub.status.idle":"2023-02-20T21:29:14.019168Z","shell.execute_reply.started":"2023-02-20T21:29:05.674994Z","shell.execute_reply":"2023-02-20T21:29:14.017944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BASE_PATH = '/kaggle/input/rsna-bcd-roi-1024x512-png-v2-dataset'\n\n# train\ndf = pd.read_csv(f'{BASE_PATH}/train.csv')\ndf['image_path'] = f'{BASE_PATH}/train_images'\\\n                    + '/' + df.patient_id.astype(str)\\\n                    + '/' + df.image_id.astype(str)\\\n                    + '.png'\nprint('Train:')\ndisplay(df.head(2))\n\n# test\ntest_df = pd.read_csv(f'{BASE_PATH}/test.csv')\ntest_df['image_path'] = f'{BASE_PATH}/test_images'\\\n                    + '/' + test_df.patient_id.astype(str)\\\n                    + '/' + test_df.image_id.astype(str)\\\n                    + '.png'\nprint('\\nTest:')\ndisplay(test_df.head(2))","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:14.020756Z","iopub.execute_input":"2023-02-20T21:29:14.021475Z","iopub.status.idle":"2023-02-20T21:29:14.318649Z","shell.execute_reply.started":"2023-02-20T21:29:14.021435Z","shell.execute_reply":"2023-02-20T21:29:14.317477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To protect the original training data frame, I saved the train file in my PC directory as file1.csv.","metadata":{}},{"cell_type":"code","source":"df.to_csv('file1.csv')","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:14.32035Z","iopub.execute_input":"2023-02-20T21:29:14.322984Z","iopub.status.idle":"2023-02-20T21:29:14.675886Z","shell.execute_reply.started":"2023-02-20T21:29:14.322941Z","shell.execute_reply":"2023-02-20T21:29:14.674861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Check the data","metadata":{}},{"cell_type":"code","source":"tf.io.gfile.exists(df.image_path.iloc[0]), tf.io.gfile.exists(test_df.image_path.iloc[0])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:14.677107Z","iopub.execute_input":"2023-02-20T21:29:14.678027Z","iopub.status.idle":"2023-02-20T21:29:14.696993Z","shell.execute_reply.started":"2023-02-20T21:29:14.677986Z","shell.execute_reply":"2023-02-20T21:29:14.695763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df2 = pd.read_csv('/kaggle/working/file1.csv')\ndf2.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:14.703398Z","iopub.execute_input":"2023-02-20T21:29:14.704424Z","iopub.status.idle":"2023-02-20T21:29:14.896765Z","shell.execute_reply.started":"2023-02-20T21:29:14.704383Z","shell.execute_reply":"2023-02-20T21:29:14.895397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport cv2\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport os\n\n# Define the input CSV file and output directory\ncsv_file = \"file1.csv\"\noutput_dir = \"processed_images\"\n\n# Create the output directory if it doesn't exist\nif not os.path.exists(output_dir):\n    os.mkdir(output_dir)\n\n# Read the CSV file\ndf = pd.read_csv(csv_file)\n\n\n#Process the images in smaller batches: Rather than processing all the images at once, you can split them into smaller batches \n# and process each batch separately. This can be done by modifying the loop that loads and resizes the images. For example, \n# you can process 1000 images at a time like this:\n\n# Process the images in smaller batches\nbatch_size = 1000\nnum_batches = len(df) // batch_size + 1\n\n# Create a new DataFrame to store the processed image paths\nprocessed_df = pd.DataFrame(columns=['image_path', 'processed_path'])\n\nfor i in range(num_batches):\n    start_idx = i * batch_size\n    end_idx = min((i+1) * batch_size, len(df))\n    processed_paths = []\n    for j, path in enumerate(df['image_path'][start_idx:end_idx]):\n        # Check if the current index is within the valid range\n        if j + start_idx >= len(df):\n            break\n\n        # Load the image\n        img = cv2.imread(path, cv2.IMREAD_GRAYSCALE)\n\n        # you can try printing the file path for each image that is \n        # being processed to check if the path is correct. You can also add\n        # a check to make sure that the image dimensions are not empty before resizing it. \n        if img is None:\n            print(\"Error loading image file:\", path)\n            continue\n\n        # Check that the image dimensions are not empty\n        if img.shape[0] == 0 or img.shape[1] == 0:\n            print(\"Error: Empty image dimensions:\", path)\n            continue\n\n        # Resize the image\n        img_resized = cv2.resize(img, (256, 256))\n\n        # Normalize the image and apply the \"viridis\" colormap\n        img_normalized = cv2.normalize(img_resized, None, alpha=0, beta=255, norm_type=cv2.NORM_MINMAX, dtype=cv2.CV_8UC1)\n        img_viridis = cv2.applyColorMap(img_normalized.astype(np.uint8), cv2.COLORMAP_VIRIDIS)\n        #colormap = plt.get_cmap('viridis')\n        #img_viridis = colormap(image_normalized)[:, :, :3]  # Remove the alpha channel\n\n        # Save the processed image as a PNG file with the original filename\n        output_filename = os.path.basename(path)[:-4] + \"_processed.png\"\n        output_path = os.path.join(output_dir, output_filename)\n        cv2.imwrite(output_path, img_viridis)\n\n        # Add the processed image path to the list\n        processed_paths.append(output_path)\n\n    # Create a new DataFrame with the processed image paths\n    batch_df = pd.DataFrame({'image_path': df['image_path'][start_idx:end_idx], 'processed_path': processed_paths})\n\n    # Append the batch DataFrame to the processed DataFrame\n    processed_df = pd.concat([processed_df, batch_df])\n\n# Merge the processed DataFrame with the original DataFrame\ndf = pd.merge(df, processed_df, on='image_path', how='left')\n\n# Save the updated DataFrame to the CSV file\ndf.to_csv(csv_file, index=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:29:14.898565Z","iopub.execute_input":"2023-02-20T21:29:14.899008Z","iopub.status.idle":"2023-02-20T21:48:46.468396Z","shell.execute_reply.started":"2023-02-20T21:29:14.898961Z","shell.execute_reply":"2023-02-20T21:48:46.465299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The photos in your CSV file are assumed to be in PNG format; this code imports them using cv2.IMREAD_UNCHANGED to keep any transparency data intact. In a further step, Matplotlib uses cv2.resize() to scale the pictures down before applying the \"viridis\" colormap. At last, it creates a new directory named \"processed_images\" to store the PNG files that result from the processing, and adds a column to the original CSV file labeled \"processed_path\" to store the locations of these files.\n\nAfter the photos are saved, the code appends their filenames to a list named \"processed_paths,\" and then inserts a new column into the original CSV file named \"processed_path\" to store the filenames. The to csv() method is then used to replace the existing CSV file with the modified one.\n\nUsing the same code as before, but with the column name changed to \"processed_path,\" you can now import the processed photos from the CSV file.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n# Select only 1 row with cancer equal to 0\nnon_cancer_df = df[df['cancer'] == 0].head(1)\n# Select only 1 row with cancer equal to 1\ncancer_df  = df[df['cancer'] == 1].head(1)\n# Select only 1 row with implant equal to 1\nimplant_df = df[df['implant'] == 1].head(1)\n\nselected_df = pd.concat([cancer_df, non_cancer_df, implant_df])\n\n# Plot the color intensity density plot for each image\nfor _, row in selected_df.iterrows():\n    # Load the image\n    img = plt.imread(row['processed_path'])\n\n    # Flatten the image to a 1D array\n    img_flattened = img.flatten()\n\n    # Create a density plot of the image color intensity values\n    sns.kdeplot(img_flattened, fill=True)\n\n# Add a legend to the plot\nplt.legend(selected_df['cancer'].astype(str))\n\n# Set the title and axis labels for the plot\nplt.title('Density Plot of Image Color Intensity')\nplt.xlabel('Color Intensity')\nplt.ylabel('Density')\n\n# Show the plot\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:49:42.439035Z","iopub.execute_input":"2023-02-20T21:49:42.43959Z","iopub.status.idle":"2023-02-20T21:49:45.509199Z","shell.execute_reply.started":"2023-02-20T21:49:42.439551Z","shell.execute_reply":"2023-02-20T21:49:45.507723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code makes a KDE plot for each category using seaborn.kdeplot, after defining a dictionary of categories and the amount of photos to choose from each category. Input for the kdeplot function is the x and y positions of the color intensity values that were previously retrieved from the selected pictures with the help of the cv2.imread command. For the KDE curves to have shading applied, True must be used for the shade option. The viridis color map is then used for the shading by setting the cmap option to viridis.\n\nFirst, using boolean indexing, I pick out one row where cancer = 0, one row where cancer = 1, and one row when implant = 1. The concat() method is then used to join these three DataFrames into a single one that we'll call \"selected df.\"\n\nThen, I loop through the rows in selected df, selecting each picture to load from the processed path column with plt.imread (). Then, we use the flatten() function to convert the picture to a 1D array, and then I use sns.kdeplot to generate a density plot of the color intensity data (). Shade is True in order to completely cover the region beneath the density curve.\n\nAfter generating a density plot for each image, I use the legend() function to label the figure. I used the astype() function to convert the cancer column values to strings, and then I used those strings as the legend labels.\n\nLastly, I use plt.title(), plt.xlabel(), and plt.ylabel() to label the title and axes of the plot. With plt.show, I display the graph ().","metadata":{}},{"cell_type":"code","source":"# Picture segmentation and color intensity value density plotting\n\nimport numpy as np\nimport pandas as pd\nimport cv2\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Define the input CSV file and output directory\ncsv_file = \"file1.csv\"\noutput_dir = \"processed_images\"\n\n# Read the CSV file\ndf = pd.read_csv(csv_file)\n\n# Select only 5 rows with cancer equal to 0\nnon_cancer_df = df[df['cancer'] == 0].head()\n\n# Select only 5 rows with cancer equal to 1\ncancer_df = df[df['cancer'] == 1].head()\n\n# Select only 5 rows with implant equal to 1\nimplant_df = df[df['implant'] == 1].head()\n\n# Concatenate the selected rows\nselected_df = pd.concat([cancer_df, non_cancer_df, implant_df])\n\n\n\nfor k, row in selected_df.iterrows():\n    # Load the original image\n    img = cv2.imread(row['image_path'])\n    \n    # Load the corresponding processed image\n    proc_img = cv2.imread(row['processed_path'])\n    \n    # Threshold the processed image to segment it\n    k, seg_img = cv2.threshold(proc_img, 200, 255, cv2.THRESH_BINARY)\n\n # Plot the original and masked images\n    fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 5))\n    ax1.imshow(proc_img)\n    ax1.set_title('Cancer: ' + str(row['cancer']==1) + ', Non-Cancer: ' + str(row['cancer']==0) + ', Implant: ' + str(row['implant']==1))\n    ax2.imshow(seg_img)\n    ax2.set_title('Threshold')\n    \n    plt.show()\n   \n   # Create a density plot of the color intensity values for the masked image\n    plt.figure(figsize=(10, 5))\n    sns.kdeplot(seg_img.ravel(), color='black', warn_singular=False, fill=True)\n    plt.title( 'Cancer: ' + str(row['cancer']==1) + ', Non-Cancer: ' + str(row['cancer']==0) + ', Implant: ' + str(row['implant']==1))\n    plt.xlabel('Color Intensity')\n    plt.ylabel('Density')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:50:05.812408Z","iopub.execute_input":"2023-02-20T21:50:05.812911Z","iopub.status.idle":"2023-02-20T21:50:29.61514Z","shell.execute_reply.started":"2023-02-20T21:50:05.812868Z","shell.execute_reply":"2023-02-20T21:50:29.613822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is necessary to construct a function that segments the picture based on the intensity values of the pixels in order to remove the pixels representing adipose tissue and implants from the chosen photos. We'll use an intensity threshold to filter out any pixels that are too weak.\n","metadata":{}},{"cell_type":"code","source":"# Segmentation and apply green filter\nimport pandas as pd\nimport cv2\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport os\n\n\n# Define the input CSV file and output directory\ncsv_file = \"file1.csv\"\noutput_dir = \"treshold_images\"\n\n# Create the output directory if it doesn't exist\nif not os.path.exists(output_dir):\n    os.mkdir(output_dir)\n\n# Read the CSV file\ndf = pd.read_csv(csv_file)\n\n# Process the images in smaller batches\nbatch_size = 1000\nnum_batches = len(df) // batch_size + 1\n\n# Create a new DataFrame to store the processed image paths\nprocessed_df = pd.DataFrame(columns=['image_path', 'processed_path', 'threshold_path'])\n\nfor i in range(num_batches):\n    start_idx = i * batch_size\n    end_idx = min((i+1) * batch_size, len(df))\n    threshold_paths = []\n    for j, path in enumerate(df['processed_path'][start_idx:end_idx]):\n        # Check if the current index is within the valid range\n        if j + start_idx >= len(df):\n            break\n\n        # Load the image\n        # Load the corresponding processed image\n        proc_img = cv2.imread(path)\n\n        # you can try printing the file path for each image that is \n        # being processed to check if the path is correct. You can also add\n        # a check to make sure that the image dimensions are not empty before resizing it. \n        if proc_img is None:\n            print(\"Error loading image file:\", path)\n            continue\n\n        # Check that the image dimensions are not empty\n        if proc_img.shape[0] == 0 or proc_img.shape[1] == 0:\n            print(\"Error: Empty image dimensions:\", path)\n            continue\n\n       \n       # Threshold the processed image to segment it\n        i,seg_img = cv2.threshold(proc_img, 200, 255, cv2.THRESH_BINARY)\n\n        # Convert the image to the HSV color space\n        hsv_image = cv2.cvtColor(seg_img, cv2.COLOR_BGR2HSV)\n\n        # Define the lower and upper boundaries for the green color in HSV space\n        lower_green = np.array([36, 25, 25])\n        upper_green = np.array([70, 255, 255])\n\n       # Create a mask to filter out the green pixels\n        mask = cv2.inRange(hsv_image, lower_green, upper_green)\n\n      # Invert the mask to keep the pixels that are not green\n        mask_inv = cv2.bitwise_not(mask)\n\n      # Apply the mask to the image to eliminate the green pixels\n        result = cv2.bitwise_and(seg_img, seg_img, mask=mask_inv)\n\n        # Save the processed image as a PNG file with the original filename\n        output_filename = os.path.basename(path)[:-4] + \"_threshold.png\"\n        output_path = os.path.join(output_dir, output_filename)\n        cv2.imwrite(output_path, result)\n\n        # Add the processed image path to the list\n        threshold_paths.append(output_path)\n\n    # Create a new DataFrame with the threshold image paths\n    batch_df = pd.DataFrame({'image_path': df['image_path'][start_idx:end_idx], 'threshold_path': threshold_paths})\n\n    # Append the batch DataFrame to the processed DataFrame\n    processed_df = pd.concat([processed_df, batch_df])\n\n# Merge the processed DataFrame with the original DataFrame\ndf = pd.merge(df, processed_df, on='image_path', how='left')\n\n# Save the updated DataFrame to the CSV file\ndf.to_csv(csv_file, index=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:50:46.034048Z","iopub.execute_input":"2023-02-20T21:50:46.034483Z","iopub.status.idle":"2023-02-20T21:53:40.718075Z","shell.execute_reply.started":"2023-02-20T21:50:46.034447Z","shell.execute_reply":"2023-02-20T21:53:40.716655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is necessary to construct a function that segments the picture based on the intensity values of the pixels in order to remove the pixels representing adipose tissue and implants from the chosen photos. I use an intensity threshold to filter out any pixels that are too weak.\n\n\nHere, I use cv2.cvtColor to transform the segmented picture from BGR to HSV color space. Finally, we make a mask using cv2.inRange to exclude the green pixels, and we set the bottom and higher borders for the green color in HSV space. For the non-green pixels, I use cv2.bitwise not to invert the mask. After that, I use cv2.bitwise and to apply the mask to the input picture, which removes the green pixels, and cv2.imshow to show both the original and modified images.","metadata":{}},{"cell_type":"code","source":"# segmented cancer image visualization after green filtering.\nimport cv2\nimport numpy as np\nimport matplotlib.pyplot as plt\n\n# Load the image\nimage = cv2.imread('/kaggle/input/image/59455684_processed_seg.png')\n\n# Convert the image to the HSV color space\nhsv_image = cv2.cvtColor(image, cv2.COLOR_BGR2HSV)\n\n# Define the lower and upper boundaries for the green color in HSV space\nlower_green = np.array([36, 25, 25])\nupper_green = np.array([70, 255, 255])\n\n# Create a mask to filter out the green pixels\nmask = cv2.inRange(hsv_image, lower_green, upper_green)\n\n# Invert the mask to keep the pixels that are not green\nmask_inv = cv2.bitwise_not(mask)\n\n# Apply the mask to the image to eliminate the green pixels\nresult = cv2.bitwise_and(image, image, mask=mask_inv)\n\n# Display the original image and the result using matplotlib\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(10, 5))\nfig.suptitle('Original Image vs Result')\nax1.imshow(cv2.cvtColor(image, cv2.COLOR_BGR2RGB))\nax1.set_title('Original Image')\nax1.axis('off')\nax2.imshow(cv2.cvtColor(result, cv2.COLOR_BGR2RGB))\nax2.set_title('Result')\nax2.axis('off')\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:58:31.677081Z","iopub.execute_input":"2023-02-20T21:58:31.67755Z","iopub.status.idle":"2023-02-20T21:58:31.859285Z","shell.execute_reply.started":"2023-02-20T21:58:31.677513Z","shell.execute_reply":"2023-02-20T21:58:31.858207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <font color='red'> Modeling </font> ","metadata":{}},{"cell_type":"markdown","source":"I use transfer learning to build a model for breast cancer classification.\n1. Get a model that has already been trained; we'll use the RESNET50 model that was trained with the ImageNet data set. The model will be loaded, but the final few layers will be discarded.\n2. To make the pre-trained model better fit our needs, we will add a few unique layers on top of it. After the initial convolutional layer, we'll add a sigmoid activation function dense layer, followed by a dropout layer, and finally a global average pooling layer.\n3. Freeze pre-trained layers: We'll freeze the pre-trained layers to prevent them from being updated during training.\n4. Compile the model: We'll use binary cross-entropy loss and the Adam optimizer.\n5. Train the model: We'll train the model on the training data and evaluate its performance on the validation data.\n6. Fine-tune the model: After training the custom layers, we can fine-tune the pre-trained layers to further improve performance. We'll unfreeze some of the top layers and continue training the model.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport tensorflow as tf\nfrom tensorflow.keras import layers, models\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import cohen_kappa_score, confusion_matrix\nimport matplotlib.pyplot as plt\nfrom sklearn.metrics import f1_score\nfrom tensorflow.keras.losses import BinaryCrossentropy\n# Vis Libs..\nfrom sklearn.manifold import TSNE\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\nplt.rcParams[\"axes.grid\"] = False","metadata":{"execution":{"iopub.status.busy":"2023-02-27T12:29:43.50165Z","iopub.execute_input":"2023-02-27T12:29:43.502145Z","iopub.status.idle":"2023-02-27T12:29:53.626282Z","shell.execute_reply.started":"2023-02-27T12:29:43.502109Z","shell.execute_reply":"2023-02-27T12:29:53.624996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the CSV file\ndf = pd.read_csv('/kaggle/working/file1.csv')\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T12:29:57.550767Z","iopub.execute_input":"2023-02-27T12:29:57.551747Z","iopub.status.idle":"2023-02-27T12:29:57.908889Z","shell.execute_reply.started":"2023-02-27T12:29:57.551697Z","shell.execute_reply":"2023-02-27T12:29:57.906803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BATCH_SIZE = 16\nIMAGE_SIZE = (224, 224)\n\n# Define a generator function to load the images in batches\ndef image_generator(df, batch_size=BATCH_SIZE, image_size=IMAGE_SIZE):\n    num_images = len(df)\n    num_batches = int(np.ceil(num_images / batch_size))\n    while True:\n        for i in range(num_batches):\n            start_index = i * batch_size\n            end_index = min((i + 1) * batch_size, num_images)\n            batch_images = []\n            for img_path in df['threshold_path'].iloc[start_index:end_index]:\n                img = tf.keras.preprocessing.image.load_img(img_path, target_size=image_size)\n                img_arr = tf.keras.preprocessing.image.img_to_array(img)\n                batch_images.append(img_arr)\n            batch_images = np.array(batch_images)\n            batch_images = batch_images / 255.0\n            batch_labels = np.array(df['cancer'].iloc[start_index:end_index])\n            yield batch_images, batch_labels","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:58:47.752368Z","iopub.execute_input":"2023-02-20T21:58:47.75282Z","iopub.status.idle":"2023-02-20T21:58:47.762207Z","shell.execute_reply.started":"2023-02-20T21:58:47.752777Z","shell.execute_reply":"2023-02-20T21:58:47.760912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code defines a generator function called image_generator that loads images in batches from a pandas dataframe.\n\nThe function takes in three arguments:\n\ndf: the dataframe that contains the file paths and labels for the images.\nbatch_size: the size of the batch to load.\nimage_size: the target size of the images.\nThe function first calculates the number of images and batches based on the batch size. It then loops through the number of batches and loads images in batches using tf.keras.preprocessing.image.load_img() and tf.keras.preprocessing.image.img_to_array(). The images are then normalized by dividing by 255.0.\n\nThe function yields a batch of images and their corresponding labels, and can be used as a generator to train a model using model.fit().","metadata":{}},{"cell_type":"code","source":"# Split the data into training and testing sets\ntrain_df, test_df = train_test_split(df, test_size=0.2, random_state=42)\nprint(train_df.shape,test_df.shape)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:58:51.064872Z","iopub.execute_input":"2023-02-20T21:58:51.065314Z","iopub.status.idle":"2023-02-20T21:58:51.101256Z","shell.execute_reply.started":"2023-02-20T21:58:51.065279Z","shell.execute_reply":"2023-02-20T21:58:51.099804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''This Function Plots a Bar plot of output Classes Distribution'''\n\ndef plot_classes(df,title):\n    df_group = pd.DataFrame(df.groupby('cancer').agg('size').reset_index())\n    df_group.columns = ['cancer','count']\n\n    sns.set(rc={'figure.figsize':(10,5)}, style = 'whitegrid')\n    sns.barplot(x = 'cancer',y='count',data = df_group,palette = \"Blues_d\")\n    plt.title('Output Class Distribution ' + str(title))\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:58:55.493773Z","iopub.execute_input":"2023-02-20T21:58:55.494329Z","iopub.status.idle":"2023-02-20T21:58:55.504806Z","shell.execute_reply.started":"2023-02-20T21:58:55.494278Z","shell.execute_reply":"2023-02-20T21:58:55.503294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_classes(train_df,\"TRAIN DATA\")","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:58:58.098185Z","iopub.execute_input":"2023-02-20T21:58:58.098606Z","iopub.status.idle":"2023-02-20T21:58:58.321817Z","shell.execute_reply.started":"2023-02-20T21:58:58.098572Z","shell.execute_reply":"2023-02-20T21:58:58.320617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_classes(test_df,\"TRAIN DATA\")","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:59:01.26199Z","iopub.execute_input":"2023-02-20T21:59:01.263223Z","iopub.status.idle":"2023-02-20T21:59:01.801019Z","shell.execute_reply.started":"2023-02-20T21:59:01.263171Z","shell.execute_reply":"2023-02-20T21:59:01.799709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create the CNN model\nmodel = models.Sequential([\n    layers.Conv2D(32, (3, 3), activation='relu', input_shape=(IMAGE_SIZE[0], IMAGE_SIZE[1], 3)),\n    layers.MaxPooling2D((2, 2)),\n    layers.Conv2D(64, (3, 3), activation='relu'),\n    layers.MaxPooling2D((2, 2)),\n    layers.Flatten(),\n    layers.Dense(64, activation='relu'),\n    layers.Dense(1, activation='sigmoid')\n])\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:59:36.426473Z","iopub.execute_input":"2023-02-20T21:59:36.426918Z","iopub.status.idle":"2023-02-20T21:59:36.651676Z","shell.execute_reply.started":"2023-02-20T21:59:36.426884Z","shell.execute_reply":"2023-02-20T21:59:36.650406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:59:40.085925Z","iopub.execute_input":"2023-02-20T21:59:40.086352Z","iopub.status.idle":"2023-02-20T21:59:40.118407Z","shell.execute_reply.started":"2023-02-20T21:59:40.086314Z","shell.execute_reply":"2023-02-20T21:59:40.117467Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is a sequential model with 5 layers, using the following layers:\n\n1. Conv2D layer with 32 filters of size 3x3, ReLU activation, and input shape of (IMAGE_SIZE[0], IMAGE_SIZE[1], 3) which corresponds to a 2D image of height IMAGE_SIZE[0] and width IMAGE_SIZE[1] with 3 color channels (RGB).\n2. MaxPooling2D layer with pool size of 2x2.\n3. Conv2D layer with 64 filters of size 3x3 and ReLU activation.\n4. MaxPooling2D layer with pool size of 2x2.\n5. Flatten layer to convert the output from the convolutional layers to a 1D array.\n6. Dense layer with 64 units and ReLU activation.\n7. Dense layer with a single output unit and sigmoid activation, which corresponds to binary classification.\n\nThis model is suitable for image classification tasks where the input images are of size IMAGE_SIZE, and the output is binary (either 0 or 1).","metadata":{}},{"cell_type":"markdown","source":"This is the summary of a convolutional neural network model with two convolutional layers, two max pooling layers, and two dense layers. The input shape is (510, 510, 3), and there are 32 filters in the first convolutional layer and 64 filters in the second convolutional layer. The dense layer has 64 units and the output layer has one unit with a sigmoid activation function, which is used for binary classification. The model has 11,963,457 total parameters, all of which are trainable.","metadata":{}},{"cell_type":"code","source":"def focal_loss(gamma=2.0, alpha=0.25):\n    # Define the focal loss function\n    def loss(y_true, y_pred):\n        # Calculate the binary cross-entropy loss\n        bce_loss = BinaryCrossentropy()(y_true, y_pred)\n        # Calculate the focal loss\n        pt = tf.where(tf.equal(y_true, 1), y_pred, 1 - y_pred)\n        focal_loss = alpha * tf.pow(1 - pt, gamma) * bce_loss\n        # Return the total loss\n        return tf.reduce_mean(focal_loss)\n    # Return the loss function\n    return loss","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:59:43.648497Z","iopub.execute_input":"2023-02-20T21:59:43.648936Z","iopub.status.idle":"2023-02-20T21:59:43.657189Z","shell.execute_reply.started":"2023-02-20T21:59:43.6489Z","shell.execute_reply":"2023-02-20T21:59:43.655793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To address class imbalance in classification problems, a sub-loss called Focal Loss was developed from Cross-Entropy. Several classification issues involve a situation where the vast majority of samples fall into one class and just a small percentage fall into the other class. Due to this disparity, the model may end up being biased in favor of the dominant class, which would have disastrous effects for the minority class.\n\nFocal Loss solves this problem by giving less weight to the simple examples that are often used for training. That is to say, the loss function gives more weight to the incorrectly categorized cases and less to the correctly identified ones. As a consequence, the model can improve its outcomes by giving greater consideration to the underrepresented group while being trained. Object recognition and segmentation tasks, in which the majority of picture pixels belong to the background class, benefit greatly from Focal Loss.","metadata":{}},{"cell_type":"code","source":"# Compile the model\nmodel.compile(optimizer='adam', loss=focal_loss(), metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:59:46.7607Z","iopub.execute_input":"2023-02-20T21:59:46.761508Z","iopub.status.idle":"2023-02-20T21:59:46.779176Z","shell.execute_reply.started":"2023-02-20T21:59:46.761468Z","shell.execute_reply":"2023-02-20T21:59:46.777683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# # Train the model\ntrain_steps = int(np.ceil(len(train_df) / BATCH_SIZE))\ntest_steps = int(np.ceil(len(test_df) / BATCH_SIZE))\ntrain_gen = image_generator(train_df)\ntest_gen = image_generator(test_df)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:59:49.467731Z","iopub.execute_input":"2023-02-20T21:59:49.468145Z","iopub.status.idle":"2023-02-20T21:59:49.474293Z","shell.execute_reply.started":"2023-02-20T21:59:49.468109Z","shell.execute_reply":"2023-02-20T21:59:49.473177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Oversample the minority class using class_weight parameter\n# Train the model on the resampled data\nclass_weights = {0: 1., 1: 10.}\nhistory= model.fit(train_gen, epochs=1, steps_per_epoch=train_steps, validation_data=test_gen, validation_steps=test_steps, class_weight=class_weights)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:48:47.150466Z","iopub.status.idle":"2023-02-20T21:48:47.151017Z","shell.execute_reply.started":"2023-02-20T21:48:47.150743Z","shell.execute_reply":"2023-02-20T21:48:47.150768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Evaluate the model\ntest_loss, test_acc = model.evaluate(test_gen, steps=test_steps)\nprint('Test accuracy:', test_acc)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:48:47.152889Z","iopub.status.idle":"2023-02-20T21:48:47.153438Z","shell.execute_reply.started":"2023-02-20T21:48:47.153159Z","shell.execute_reply":"2023-02-20T21:48:47.153183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import roc_curve, auc\n# Get the true labels for the test set\ny_true = test_df['cancer']\ny_pred = model.predict(test_gen, steps=test_steps) > 0.5\n","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:48:47.154851Z","iopub.status.idle":"2023-02-20T21:48:47.16384Z","shell.execute_reply.started":"2023-02-20T21:48:47.163352Z","shell.execute_reply":"2023-02-20T21:48:47.16342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = y_pred.astype(int)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the ROC curve and AUC score\nfpr, tpr, thresholds = roc_curve(y_true, y_pred)\nroc_auc = auc(fpr, tpr)\n\n# Plot the ROC curve\nplt.figure()\nplt.plot(fpr, tpr, color='darkorange', lw=2, label='ROC curve (area = %0.2f)' % roc_auc)\nplt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--')\nplt.xlim([0.0, 1.0])\nplt.ylim([0.0, 1.05])\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive Rate')\nplt.title('Receiver operating characteristic')\nplt.legend(loc=\"lower right\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:48:47.165584Z","iopub.status.idle":"2023-02-20T21:48:47.166193Z","shell.execute_reply.started":"2023-02-20T21:48:47.16591Z","shell.execute_reply":"2023-02-20T21:48:47.165936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"kappa = cohen_kappa_score(y_true, y_pred)\nprint('Cohen Kappa score:', kappa)","metadata":{"execution":{"iopub.status.busy":"2023-02-20T21:48:47.167865Z","iopub.status.idle":"2023-02-20T21:48:47.168431Z","shell.execute_reply.started":"2023-02-20T21:48:47.168149Z","shell.execute_reply":"2023-02-20T21:48:47.168175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If your model has a high accuracy score but a low Cohen Kappa score on the training data, it is likely that it is incorrectly classifying the vast majority of the samples in the training set and is hence failing to capture the underlying variability in the data. This may occur if the data is skewed, if the model is overfitting to the training data, or if it is not sophisticated enough to learn the characteristics in the data.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\n\n# Convert y_pred to a DataFrame\npred_df = pd.DataFrame(y_pred, columns=['cancer'])\n\n# Add the image names as a column\npred_df['image_id'] = test_df['image_id']\n\n# Save to a CSV file\npred_df.to_csv('cancer_predictions.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-02-27T12:29:33.891091Z","iopub.execute_input":"2023-02-27T12:29:33.891655Z","iopub.status.idle":"2023-02-27T12:29:33.917147Z","shell.execute_reply.started":"2023-02-27T12:29:33.891611Z","shell.execute_reply":"2023-02-27T12:29:33.915066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load patient information from file1.csv\npatient_df = pd.read_csv('file1.csv')\n\n# Load cancer predictions from cancer_predictions.csv\npred_df = pd.read_csv('cancer_predictions.csv')\n\n# Combine patient_id and laterality to create the prediction_id\npred_df['prediction_id'] = patient_df['patient_id'].astype(str) + '_' + patient_df['laterality']\n\n# Extract the prediction probability as a column named \"cancer\"\npred_df['cancer'] = pred_df['cancer']\n\n# Drop the original index from the predictions DataFrame\npred_df = pred_df.reset_index(drop=True)\n\n# Select only the columns needed for the submission file\nsubmission_df = pred_df[['prediction_id', 'cancer']]\n\n# Save the submission file\nsubmission_df.to_csv('submission_df.csv', index=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-02-23T14:33:19.803287Z","iopub.execute_input":"2023-02-23T14:33:19.803648Z","iopub.status.idle":"2023-02-23T14:33:19.833899Z","shell.execute_reply.started":"2023-02-23T14:33:19.803618Z","shell.execute_reply":"2023-02-23T14:33:19.83264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission_df = pd.read_csv('submission_df.csv')\nsubmission = submission_df.head(2)\nsubmission.to_csv('submission.csv', index=False)\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2023-02-27T12:29:30.824437Z","iopub.execute_input":"2023-02-27T12:29:30.825043Z","iopub.status.idle":"2023-02-27T12:29:30.926017Z","shell.execute_reply.started":"2023-02-27T12:29:30.824937Z","shell.execute_reply":"2023-02-27T12:29:30.924471Z"},"trusted":true},"execution_count":null,"outputs":[]}]}