{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"For fun, going to ignore the images here and just do EDA on tabular data.  \nWhether this data will be relevant for test is questionable and some of these insights may be purely coincidental.\n\nThere are some cool auto eda libraries, check them out - https://medium.com/geekculture/10-automated-eda-libraries-at-one-place-ea5d4c162bbb\nThey are actually very useful and I recommend using them, however, in a notebook like this they are a bit noisy.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import pandas as pd\ntrain = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/train.csv\")\norigtrain = train.copy()\ntest = pd.read_csv(\"/kaggle/input/rsna-breast-cancer-detection/test.csv\")\ntrain.info()\n#Drop the patient/image id\ncols = ['site_id', 'laterality', 'view', 'age', 'cancer', 'biopsy', 'BIRADS', 'implant',\n        'density', 'machine_id', 'difficult_negative_case', 'invasive', 'patient_id']\ntrain = train[cols]\ntrain['laterality'] = train['laterality'].apply(lambda x:1 if x == 'L' else 0)\ntmp = train[['view', 'density']]\ntrain = pd.get_dummies(train)\ntrain['view'] = tmp['view']\ntrain['density'] = tmp['density']","metadata":{"execution":{"iopub.status.busy":"2022-12-07T20:16:55.585771Z","iopub.execute_input":"2022-12-07T20:16:55.586161Z","iopub.status.idle":"2022-12-07T20:16:55.727792Z","shell.execute_reply.started":"2022-12-07T20:16:55.58613Z","shell.execute_reply":"2022-12-07T20:16:55.726701Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def pfbeta(labels, predictions, beta = 1):\n    y_true_count = 0\n    ctp = 0\n    cfp = 0\n\n    for idx in range(len(labels)):\n        prediction = min(max(predictions[idx], 0), 1)\n        if (labels[idx]):\n            y_true_count += 1\n            ctp += prediction\n        else:\n            cfp += prediction\n\n    beta_squared = beta * beta\n    c_precision = ctp / (ctp + cfp)\n    c_recall = ctp / y_true_count\n    if (c_precision > 0 and c_recall > 0):\n        result = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n        return result\n    else:\n        return 0\n\nfor k in train['machine_id'].unique():\n    print(k, pfbeta(train['cancer'], [0 if x != k else 1 for x in train['machine_id']]))\n    \nfor k in train['site_id'].unique():\n    print(k, pfbeta(train['cancer'], [0 if x != k else 1 for x in train['site_id']]))","metadata":{"execution":{"iopub.status.busy":"2022-12-07T19:26:37.755174Z","iopub.execute_input":"2022-12-07T19:26:37.755667Z","iopub.status.idle":"2022-12-07T19:26:40.668604Z","shell.execute_reply.started":"2022-12-07T19:26:37.755638Z","shell.execute_reply":"2022-12-07T19:26:40.667683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"pfbeta scores per machine if you set them to 1 in train.","metadata":{}},{"cell_type":"code","source":"cpsL = origtrain.query(\"cancer == 1 and laterality == 'L'\")['patient_id'].unique()\ncpsR = origtrain.query(\"cancer == 1 and laterality == 'R'\")['patient_id'].unique()\ncL,cR = 0,0\nfor i,r in origtrain.iterrows():\n    if r['cancer']==0 and r['laterality']=='L' and r['patient_id'] in cpsL:\n        cL = cL + 1\n    if r['cancer']==0 and r['laterality']=='R' and r['patient_id'] in cpsR:\n        cR = cR + 1\nlen(cpsL), len(cpsR), cL,cR, len([x for x in cpsL if x not in cpsR]),len([x for x in cpsR if x not in cpsL])\ntrain = train.drop_duplicates([\"patient_id\", \"laterality\"])\nprint(\"left site 1,\", len(train.query(\"laterality == 1 and site_id == 1\")))\nprint(\"right site 1,\", len(train.query(\"laterality == 0 and site_id == 1\")))\nprint(\"left site 2,\", len(train.query(\"laterality == 1 and site_id == 2\")))\nprint(\"right site 2,\", len(train.query(\"laterality == 0 and site_id == 2\")))","metadata":{"execution":{"iopub.status.busy":"2022-12-07T20:31:30.257966Z","iopub.execute_input":"2022-12-07T20:31:30.258615Z","iopub.status.idle":"2022-12-07T20:31:34.641337Z","shell.execute_reply.started":"2022-12-07T20:31:30.258557Z","shell.execute_reply":"2022-12-07T20:31:34.640132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just verifying that all patients with cancer are consistently identified, ie: you don't get a view that says cancern and another view that says no cancer for the same laterality.  This is true.  Cancer in one laterality rarely means cancer in the other, though less so for the right laterality.\n\nSince they are consistently identified, we're going to drop duplicates here.  Note that this only checks for cancer, the other columns need to be investigate for abnormal distributions.","metadata":{}},{"cell_type":"code","source":"gb = train.groupby(\"machine_id\").mean().sort_values(\"cancer\", ascending = False)\ndisplay(gb)\ngb = train.groupby(\"machine_id\").count()['age']\ndisplay(gb)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-12-07T20:26:29.650279Z","iopub.execute_input":"2022-12-07T20:26:29.651415Z","iopub.status.idle":"2022-12-07T20:26:29.700279Z","shell.execute_reply.started":"2022-12-07T20:26:29.651358Z","shell.execute_reply":"2022-12-07T20:26:29.699095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Machine 190 seems to be getting the highest cancer rate but the sample space is very low, so difficult to conclude from that.  Machine 49 \"AT\" views which is a sign that the staff is worried and is getting more views done they don't normally do (see below, much higher indicidence of cancer with AT view).\n\n\nIt is interesting to note machines 49/170 have the highest biopsy rates.","metadata":{}},{"cell_type":"code","source":"gb_age = train.groupby(\"age\").mean().sort_values(\"cancer\", ascending = False).reset_index()\nprint(\"Age/Cancer correlation\", gb_age['age'].corr(gb_age['cancer']))\ndisplay(gb_age)","metadata":{"execution":{"iopub.status.busy":"2022-12-07T20:28:03.24994Z","iopub.execute_input":"2022-12-07T20:28:03.250372Z","iopub.status.idle":"2022-12-07T20:28:03.296261Z","shell.execute_reply.started":"2022-12-07T20:28:03.250337Z","shell.execute_reply":"2022-12-07T20:28:03.295134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Age is obviously a factor in cancer diagnosis.  Fairly well understood.","metadata":{}},{"cell_type":"code","source":"gb_implant = train.groupby([\"implant\",\"site_id\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb_implant)\nprint(len(train.query(\"implant == 1 and cancer == 1\")))\nprint(len(train.query(\"cancer == 1\")))","metadata":{"execution":{"iopub.status.busy":"2022-12-07T20:28:13.410413Z","iopub.execute_input":"2022-12-07T20:28:13.410796Z","iopub.status.idle":"2022-12-07T20:28:13.454878Z","shell.execute_reply.started":"2022-12-07T20:28:13.410765Z","shell.execute_reply":"2022-12-07T20:28:13.453842Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only site_id 1 lists implants.\nAlmost 3 times more likely to get cancer if you don't have an implant.  Interesting.\nDo note the age and the small sample space.  \n\nThere is also this:  Implants may slightly risk the incidence of cancer, but people who get them\n    might be less prone to getting cancer.\n\n\"A few studies in the table below found a decreased risk of breast cancer among women with \nimplants. However, this is likely due to traits of women who tend to choose breast implants \n(such as being lean), rather than the implants themselves [2].\"\n\nhttps://www.komen.org/breast-cancer/facts-statistics/research-studies/topics/breast-implants-and-breast-cancer-risk/\n\n","metadata":{}},{"cell_type":"code","source":"gb_site = train.groupby(\"site_id\").mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb_site)\ngb_site = train.groupby(\"site_id\").count()['age']\ndisplay(gb_site)\n","metadata":{"execution":{"iopub.status.busy":"2022-12-07T20:28:31.7916Z","iopub.execute_input":"2022-12-07T20:28:31.792672Z","iopub.status.idle":"2022-12-07T20:28:31.83575Z","shell.execute_reply.started":"2022-12-07T20:28:31.792628Z","shell.execute_reply":"2022-12-07T20:28:31.834588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Interesting, lower age but higher cancer indicidence at site 1.  There could be a lot of reasons for that. It could just be the population being sampled, or it could be related to the fact there seems to be a lot of missing data for site_id 2.  \n\nFor example, you were more likely to get a biopsy at site 1.  There are no \"AT\" views (high incidence of cancer, see below) at site 2.\n\nFortunately, invasive at site 2 seems to be also lower. \n\nI'm going to include group by site in the following eda.","metadata":{}},{"cell_type":"code","source":"gb_lat = train.groupby([\"laterality\", \"site_id\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb_lat)\ngb_lat = train.groupby([\"laterality\", \"site_id\"]).count()['age']\ndisplay(gb_lat)\ngb_lat = origtrain.drop_duplicates(['patient_id', 'laterality']).groupby([\"laterality\", \"site_id\", 'cancer']).count()['age']\ndisplay(gb_lat)","metadata":{"execution":{"iopub.status.busy":"2022-12-07T20:38:38.576408Z","iopub.execute_input":"2022-12-07T20:38:38.576789Z","iopub.status.idle":"2022-12-07T20:38:38.654305Z","shell.execute_reply.started":"2022-12-07T20:38:38.576759Z","shell.execute_reply":"2022-12-07T20:38:38.653056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some studies have shown that the left breast is 5%-10% more likely to develop cancer than the right breast.  Interesting that this doesn't register very well at site 2.  \n\nThe difference in counts between L and R at site 1 is surprising.\n\nNote that the data has been enriched, so it could be a case of sufficient examples of cancer were provided for each laterality at site 2.\n\n> However, the RSNA team invested a lot of effort in enriching the competition dataset for cancer cases to ensure there are enough of them to use for modeling. From memory, the cancer rate in the general population is closer to 0.1% (@vaillant might have a more accurate number). Since the false negatives weren't enriched there's a reasonable chance that there's literally only one in the entire dataset.\n\nhttps://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369262#2056107","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"gb = train.groupby([\"view\", \"site_id\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)\ndisplay(train.groupby([\"view\"]).count()['cancer'])","metadata":{"execution":{"iopub.status.busy":"2022-12-07T20:59:27.406453Z","iopub.execute_input":"2022-12-07T20:59:27.406865Z","iopub.status.idle":"2022-12-07T20:59:27.462369Z","shell.execute_reply.started":"2022-12-07T20:59:27.40683Z","shell.execute_reply":"2022-12-07T20:59:27.461289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"AT definitely has very notable incidence, but the count is very low. \n\nI believe this is the auxillary tail view and is done when a lesion is seen in the breast. https://radiopaedia.org/articles/axillary-view?lang=us\n\nInteresting to note there are no AT view at site 2 which may explain some discrepancies.\n\nThe rest of the data isn't in test, but still interesting, imho.","metadata":{}},{"cell_type":"code","source":"gb_biopsy = train.groupby([\"biopsy\", \"site_id\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb_biopsy)\ngb_biopsy = train.groupby([\"biopsy\", \"site_id\"]).count()['age']\ndisplay(gb_biopsy)\n","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:00:04.028466Z","iopub.execute_input":"2022-12-07T21:00:04.028868Z","iopub.status.idle":"2022-12-07T21:00:04.079474Z","shell.execute_reply.started":"2022-12-07T21:00:04.028835Z","shell.execute_reply":"2022-12-07T21:00:04.078341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The cancer indicidence is fairly obvious, though it's important to mention just because there is\nzero for cancer doesn't mean the patient didn't have cancer, only that none was detected. \n\nAs mentioned above, biopsies happened more frequently at site 1.   ","metadata":{}},{"cell_type":"code","source":"import numpy as np\nprint(\"site one NAN birad\", len(train[np.isnan(train['BIRADS']) & (train['site_id'] == 1)]))\nprint(\"site two NAN birad\", len(train[np.isnan(train['BIRADS']) & (train['site_id'] == 2)]))\ngb = train.groupby([\"BIRADS\", \"difficult_negative_case\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)\ngb = train.groupby([\"BIRADS\", \"difficult_negative_case\"]).count()['age']\ndisplay(gb)\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]))\ndisplay(origtrain.query(\"BIRADS == 0 and difficult_negative_case == False\").drop_duplicates([\"patient_id\", \"laterality\"]).mean())","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:01:39.331568Z","iopub.execute_input":"2022-12-07T21:01:39.331967Z","iopub.status.idle":"2022-12-07T21:01:39.432819Z","shell.execute_reply.started":"2022-12-07T21:01:39.331935Z","shell.execute_reply":"2022-12-07T21:01:39.431471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Birads 1 and 2 mean the breasts were rated as negative or normal, so these stats make sense.  Site 2 seems to rarely do a BIRAD score.   Note that when BIRADS is 0 and difficult negative case is False, cancer is 1.","metadata":{}},{"cell_type":"code","source":"gb = train.groupby([\"density\", \"site_id\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)\ngb = train.groupby([\"density\", \"site_id\"]).count()['age']\ndisplay(gb)","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:02:03.136253Z","iopub.execute_input":"2022-12-07T21:02:03.136668Z","iopub.status.idle":"2022-12-07T21:02:03.191729Z","shell.execute_reply.started":"2022-12-07T21:02:03.136637Z","shell.execute_reply":"2022-12-07T21:02:03.190664Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only site 1 reports density.  Not sure why this isn't in the test data.  \nMaybe it can be inferred from the images somehow?  Note the rising incidence of cancer as we get more dense, and yet it drops for C and than significantly for density D. \n\nFrom the data description - \"Extremely dense tissue can make diagnosis more difficult\"\n\nHowever, studies have shown that people with the highest density are 4 to 6 times more likely to get breast cancer.\n\n\nThe age differences on density are interesting, too, which may explain the cancer incidence.  I suppose as you get older your breasts get less dense.\n","metadata":{}},{"cell_type":"code","source":"gb = train.groupby([\"difficult_negative_case\", \"site_id\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)\ngb = train.groupby([\"difficult_negative_case\", \"site_id\"]).count()['age']\ndisplay(gb)","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:02:15.907547Z","iopub.execute_input":"2022-12-07T21:02:15.907964Z","iopub.status.idle":"2022-12-07T21:02:15.96178Z","shell.execute_reply.started":"2022-12-07T21:02:15.90793Z","shell.execute_reply":"2022-12-07T21:02:15.960771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Site 2 is more likely to report a difficult negative case in the left breast.  \n\nI'm a bit puzzled, even alarmed, by the low biopsy rate for difficult cases in site 1 versus site 2.\n\n","metadata":{}},{"cell_type":"code","source":"gb = train.groupby([\"invasive\", \"site_id\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)\ngb = train.groupby([\"invasive\", \"site_id\"]).count()['age']\ndisplay(gb)","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:03:05.805059Z","iopub.execute_input":"2022-12-07T21:03:05.805541Z","iopub.status.idle":"2022-12-07T21:03:05.857487Z","shell.execute_reply.started":"2022-12-07T21:03:05.805491Z","shell.execute_reply":"2022-12-07T21:03:05.856291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Invasive is quite high for desnity B/C.  This could be an issue of sample space, but that explanation would be pushing it. ","metadata":{}},{"cell_type":"code","source":"import numpy as np\ntrainc = train.query(\"site_id == 1\").copy()\ntrainc['sdage'] = [0 if np.isnan(x) else int(x/10) for x in trainc['age'] if np.nan]\ngb = trainc.groupby([\"density_C\", \"sdage\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:03:13.391185Z","iopub.execute_input":"2022-12-07T21:03:13.39162Z","iopub.status.idle":"2022-12-07T21:03:13.47158Z","shell.execute_reply.started":"2022-12-07T21:03:13.391584Z","shell.execute_reply":"2022-12-07T21:03:13.470252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As you can see Density_C is a very important indicator.  Use it to train your models to determine density as its often a signal for cancer. This is also backed up extensively by research. Density_D should be an even bigger signal, unfortunately has too low of sample space and doesn't seem to work very well.  Plus it's very hard to determine cancer even if its there, so we may be dealing with some false negatives.\n\n\n","metadata":{}},{"cell_type":"code","source":"import numpy as np\ntrainc = train.query(\"site_id == 1\").copy()\ntrainc['sdage'] = [0 if np.isnan(x) else int(x/10) for x in trainc['age'] if np.nan]\ngb = trainc.groupby([\"density_B\", \"sdage\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:03:44.783345Z","iopub.execute_input":"2022-12-07T21:03:44.784691Z","iopub.status.idle":"2022-12-07T21:03:44.865789Z","shell.execute_reply.started":"2022-12-07T21:03:44.784642Z","shell.execute_reply":"2022-12-07T21:03:44.864591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"B as well has lower incidence, the age 30-39 bracket could be a statistical outlier.","metadata":{}},{"cell_type":"code","source":"import numpy as np\ntrainc = train.query(\"site_id == 1\").copy()\ntrainc['sdage'] = [0 if np.isnan(x) else int(x/10) for x in trainc['age'] if np.nan]\ngb = trainc.groupby([\"density_A\", \"sdage\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)\n","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:05:04.96288Z","iopub.execute_input":"2022-12-07T21:05:04.963831Z","iopub.status.idle":"2022-12-07T21:05:05.041446Z","shell.execute_reply.started":"2022-12-07T21:05:04.963789Z","shell.execute_reply":"2022-12-07T21:05:05.040615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Density_A is also a signal, but a negative one for cancer.","metadata":{}},{"cell_type":"code","source":"import numpy as np\ntrainc = train.query(\"site_id == 1\").copy()\ntrainc['sdage'] = [0 if np.isnan(x) else int(x/10) for x in trainc['age'] if np.nan]\ngb = trainc.groupby([\"density_D\", \"sdage\"]).mean().sort_values(\"cancer\", ascending = False).reset_index()\ndisplay(gb)\nprint(\"D count\", len(trainc.query(\"density_D == 1\")))","metadata":{"execution":{"iopub.status.busy":"2022-12-07T21:05:23.588293Z","iopub.execute_input":"2022-12-07T21:05:23.588712Z","iopub.status.idle":"2022-12-07T21:05:23.677472Z","shell.execute_reply.started":"2022-12-07T21:05:23.588681Z","shell.execute_reply":"2022-12-07T21:05:23.676514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Density D is a bit awkward because of the low sample space.  Also, it becomes harder to predict cancer as explained in the data dictionary so we may be dealing with false negatives.  Note the age 50-59 bracket however.","metadata":{}},{"cell_type":"markdown","source":"That's it for now.  Leave a comment if you think I missed something or reported something incorrectly. \n\nAlso, upvotes are always appreciated :)","metadata":{}}]}