{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Key Summary","metadata":{}},{"cell_type":"markdown","source":"1. Sites 1 and 2 might need different models as they treat Implants differently - we do not know if site 2 never has any implant. Alternatively, we could have smaller model of implant and adjust blending ratio depending on whether implant is present or not.\n2. There are a few different types of machines used. One strategy could be to have blending of one general model and then one focussed model of the specific machine type. Different machines are used in the different sites , so again having just separate models for the 2 sites can serve similar purpose.\n3. Does it make sense to start with BIRADS instead of direct classification as 0 and 1 for cancer? \n    - BIRADS score of 1 and 2 should be quite straight forward \n    - BIRADS score of 0 can use a blend of dedicated model on 0's and an overall model covering 0,1,and 2.\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-07T14:25:12.441082Z","iopub.execute_input":"2022-12-07T14:25:12.441588Z","iopub.status.idle":"2022-12-07T14:25:13.602874Z","shell.execute_reply.started":"2022-12-07T14:25:12.441484Z","shell.execute_reply":"2022-12-07T14:25:13.601658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.read_csv('/kaggle/input/1-basic-eda/basic_df.csv')","metadata":{"execution":{"iopub.status.busy":"2022-12-07T14:25:13.604933Z","iopub.execute_input":"2022-12-07T14:25:13.605384Z","iopub.status.idle":"2022-12-07T14:25:14.075163Z","shell.execute_reply.started":"2022-12-07T14:25:13.605352Z","shell.execute_reply":"2022-12-07T14:25:14.073869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.drop(columns=['Unnamed: 0'],inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-12-07T14:25:14.07666Z","iopub.execute_input":"2022-12-07T14:25:14.076988Z","iopub.status.idle":"2022-12-07T14:25:14.103098Z","shell.execute_reply.started":"2022-12-07T14:25:14.07696Z","shell.execute_reply":"2022-12-07T14:25:14.10201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.columns","metadata":{"execution":{"iopub.status.busy":"2022-12-07T14:25:14.105979Z","iopub.execute_input":"2022-12-07T14:25:14.108676Z","iopub.status.idle":"2022-12-07T14:25:14.120365Z","shell.execute_reply.started":"2022-12-07T14:25:14.108625Z","shell.execute_reply":"2022-12-07T14:25:14.119421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,ax = plt.subplots(figsize=(16,4))\ndatadf = df.groupby(['site_id']).agg(patients_with_implants=(\"implant\", 'sum')).reset_index()\ndatadf['site_id'] = datadf['site_id'].astype('category')\nsns.barplot(data=datadf,\n            x='patients_with_implants',y='site_id',hue='site_id')\nplt.legend(loc='lower right');","metadata":{"execution":{"iopub.status.busy":"2022-12-07T14:25:14.122659Z","iopub.execute_input":"2022-12-07T14:25:14.123862Z","iopub.status.idle":"2022-12-07T14:25:14.418784Z","shell.execute_reply.started":"2022-12-07T14:25:14.123814Z","shell.execute_reply":"2022-12-07T14:25:14.417278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig,ax = plt.subplots(figsize=(16,4))\ndatadf = df.groupby(['site_id','machine_id']).agg(count_machine=(\"machine_id\", 'count')).reset_index()\ndatadf['site_id'] = datadf['site_id'].astype('category')\ndatadf['machine_id'] = datadf['machine_id'].astype('category')\nsns.barplot(data=datadf,\n            x='count_machine',y='machine_id',hue='site_id')\nplt.legend(loc='lower right');","metadata":{"execution":{"iopub.status.busy":"2022-12-07T14:25:14.420757Z","iopub.execute_input":"2022-12-07T14:25:14.421629Z","iopub.status.idle":"2022-12-07T14:25:14.76423Z","shell.execute_reply.started":"2022-12-07T14:25:14.42158Z","shell.execute_reply":"2022-12-07T14:25:14.763061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.groupby(['laterality','BIRADS','biopsy','cancer']).agg(count_machine=(\"site_id\", 'count'))","metadata":{"execution":{"iopub.status.busy":"2022-12-07T14:58:23.608513Z","iopub.execute_input":"2022-12-07T14:58:23.608933Z","iopub.status.idle":"2022-12-07T14:58:23.641741Z","shell.execute_reply.started":"2022-12-07T14:58:23.608901Z","shell.execute_reply":"2022-12-07T14:58:23.640255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is some very clear correlation here w.r.t BIRADS of 1 and 2 --> indicates negative or benign\n- 100% of cases with BIRADS 1 is negative so cancer is 0\n- 100% of cases with BIRADS 2 is benign so again, cancer is 0\n- Less than 10% of cases with BIRADS of 0 is proven cancer..so 90% of prelim imaging where doctor suggests additional imaging could be one key area \n- On the other hand, only around 30% of BIRADS of 0 is actually processed for biopsy - and then a further 25-30% of those is positive for cancer","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}