{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":70367,"databundleVersionId":9188054,"sourceType":"competition"}],"dockerImageVersionId":30746,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Intro\n\nHi everyone, this is my first kaggle notebook post. In this notebook I will be sharing my attempt at feature engineering as outlined in the abstract.\n\nIn this notebook I achieve a decent correlation of **R2 = 0.945** between my engineered features and w1_target in the training data; **this is correlation between target and features only, no model has been trained**.\n\nIt is quite possible that I've not understood some conventions be it through referencing other peoples work or in the way I've constructed my notebook so will be very open to any feedback!\n\nI also apologise as I do not have too much time to work on this project so have not read too many notebooks of other people. This work will be sharing my attempts at automatic feature detection following on from the great notebooks of GORDONYIP (https://www.kaggle.com/code/aaronjday/update-calibrating-and-binning-astronomical-data) and AMBROSM (https://www.kaggle.com/code/ambrosm/adc24-intro-training).\n\nAbstract:\nAs has been explained in the competition description, each planet has a significant amount of data over the time, space, and spectral domain. As was discussed by AMBROSM in the above linked notebook, the ratio of absorbed photons during vs outside of transit should is highly correlated with absorption spectra results. In this notebook I largely follow the preprocessing advised by GORDONYIP (e.g., binning data, analgue to digital conversion) and AMBROSM (identify time steps in and outside of transit, sum over wavelengths and spatial regimes and obtain ratio of absorbed photons). I will only be using the AIRS-CH0 data. I attempt to take AMBROSMs notebook slightly forward by using a 1D convolution to identify drops in absorption to automatically identify when the planet is in transit. I then use the data outside of transit to detrend the data by fitting a polynomial and subtracting from the full time data. This method of feature engineering seems to offer a slightly improved correlation between said features and the absorption at the first wavelength.\n\nI really hope people find this method of detecting transit times via 1D convolutions and using the out-of-transit data to detrend the full time series. I think an interesting avenue for future work will be to not sum over the spectral dimension and see how the photon counts for each wavelength correlate to the spectra at that wavelength (providing the noise is not too bad)!\n\n**I have left a part of the notebook ([7]) below so you can try the method of feature engineering for any of the planets in the training data.**\n\n**[EDIT AFTER UPLOAD] Had a quick look after uploading and saw a different solution to the same problem by Pascal Pfeiffer who also did it for the FGS1 data as well! (https://www.kaggle.com/code/ilu000/ariel24-find-transit-zones-full-code)**\n\nThanks!","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nimg_A = mpimg.imread('/kaggle/working/Feature_Engineering_Method.jpg')\nimg_B = mpimg.imread('/kaggle/working/Feature_Engineering_Result.jpg')\nimg_C = mpimg.imread('/kaggle/working/Correlation.jpg')\n\n# display images\nfig, ax = plt.subplots(1,3, figsize=(18,18))\nax[0].imshow(img_A)\nax[1].imshow(img_B)\nax[2].imshow(img_C)\nax[0].axis('off')\nax[1].axis('off')\nax[2].axis('off')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-08-21T21:22:58.412207Z","iopub.execute_input":"2024-08-21T21:22:58.412675Z","iopub.status.idle":"2024-08-21T21:22:59.070384Z","shell.execute_reply.started":"2024-08-21T21:22:58.412631Z","shell.execute_reply":"2024-08-21T21:22:59.068722Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Imports","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport glob\nimport os","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:45.279275Z","iopub.execute_input":"2024-08-21T20:28:45.279656Z","iopub.status.idle":"2024-08-21T20:28:45.708337Z","shell.execute_reply.started":"2024-08-21T20:28:45.279622Z","shell.execute_reply":"2024-08-21T20:28:45.707047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_adc_info = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/train_adc_info.csv')\ntrain_labels = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:45.709707Z","iopub.execute_input":"2024-08-21T20:28:45.710236Z","iopub.status.idle":"2024-08-21T20:28:45.792229Z","shell.execute_reply.started":"2024-08-21T20:28:45.710197Z","shell.execute_reply":"2024-08-21T20:28:45.790939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploring Data","metadata":{}},{"cell_type":"code","source":"%%writefile planet_data_airs.py\n\nimport os\nimport numpy as np\nimport pandas as pd\n\nPATH_FOLDER = '/kaggle/input/ariel-data-challenge-2024/'\nNUMBER_OF_FRAMES = 11250\nNUMBER_OF_SPATIAL_POINTS = 32\nNUMBER_OF_WAVELENGTH_POINTS = 356\nLOWER_WAVELENGTH_INDEX_CUTOFF = 39\nUPPER_WAVELENGTH_INDEX_CUTOFF = 321\nNUMBER_OF_WAVELENGTH_POINTS_FILTERED = UPPER_WAVELENGTH_INDEX_CUTOFF - LOWER_WAVELENGTH_INDEX_CUTOFF\n\nTRAIN_ADC_INFO = pd.read_csv(os.path.join(PATH_FOLDER, 'train_adc_info.csv')).set_index('planet_id')\nAXIS_INFO = pd.read_parquet('/kaggle/input/ariel-data-challenge-2024/axis_info.parquet')\nDT_AIRS  = AXIS_INFO['AIRS-CH0-integration_time'].dropna().values\n\nclass PlanetDataAIRSCH0():\n    def __init__(self,\n                 planet_index: str):\n        # __init__ loads in all planet-specific information for AIRS-CH0 observations\n        self.planet_index = planet_index\n        \n        if 'TRAIN_ADC_INFO' not in globals():\n            raise Exception(\"Load in dataframe TRAIN_ADC_INFO to construct class!\")\n        \n        if 'DT_AIRS' not in globals():\n            raise Exception(\"Load DT_AIRS to construct class!\")\n        \n        # Gather raw data\n        raw_signal_path = fr'/kaggle/input/ariel-data-challenge-2024/train/{self.planet_index}/AIRS-CH0_signal.parquet'\n        raw_signal_df = pd.read_parquet(raw_signal_path)\n        raw_signal_data = raw_signal_df.to_numpy().reshape(raw_signal_df.shape[0], NUMBER_OF_SPATIAL_POINTS, NUMBER_OF_WAVELENGTH_POINTS)\n        raw_signal_data = raw_signal_data.astype(np.float64)\n        self.raw_signal_data = raw_signal_data[:,:,LOWER_WAVELENGTH_INDEX_CUTOFF:UPPER_WAVELENGTH_INDEX_CUTOFF]\n        if self.raw_signal_data.shape != (NUMBER_OF_FRAMES, NUMBER_OF_SPATIAL_POINTS, NUMBER_OF_WAVELENGTH_POINTS_FILTERED):\n            raise Exception(\"Incorrect shape for self.signal_data!\")\n        \n        self.processed_signal_data = self.raw_signal_data.copy()\n        self.list_of_applied_transformations = []\n        self.is_cds_valid = False  # CDS has not yet been calculated\n        \n        # Fetch gain and offset from TRAIN_ADC_INFO dataframe\n        self._gain = TRAIN_ADC_INFO['AIRS-CH0_adc_gain'].loc[self.planet_index]\n        self._offset = TRAIN_ADC_INFO['AIRS-CH0_adc_offset'].loc[self.planet_index]\n\n    def apply_analogue_to_digital_conversion(self):\n        # If CDS was caluclated from processed signal data, it is now invalid\n        self.is_cds_valid = False\n        # Method from cell In [5] of GORDONYIP Notebook\n        # Applies conversion inplace\n        self.processed_signal_data /= self._gain\n        self.processed_signal_data += self._offset\n        self.list_of_applied_transformations.append(\"analogue_to_digital_conversion\")\n    \n    def evaluate_cds_of_processed_signal_data(self):\n        # Function from cell In [9] of GORDONYIP Notebook\n        # Edited slightly to apply to individual result of shape (a, b, c) as opposed to array of results (n, a, b, c)\n        # Note that this\n        self.cds_of_processed_signal_data = self.processed_signal_data[1::2,:,:] - self.processed_signal_data[::2,:,:]\n        # CDS is currently valid\n        self.is_cds_valid = True\n    \n    def get_cds_of_processed_signal_data_after_binning(self,\n                                                       binning: int):\n        # Function from cell In [10] of GORDONYIP Notebook\n        # Edited slightly to apply to individual result of shape (a, b, c) as opposed to array of results (n, a, b, c)\n        if self.is_cds_valid is False:\n            raise Exception(\"CDS has either not been calculated or does not match processed signal data!\")\n        \n        cds_transposed = self.cds_of_processed_signal_data.transpose(0,2,1)\n        if binning > 1:\n            cds_binned = np.zeros((cds_transposed.shape[0]//binning, cds_transposed.shape[1], cds_transposed.shape[2]))\n            for i in range(cds_transposed.shape[0]//binning):\n                cds_binned[i,:,:] = np.sum(cds_transposed[i*binning:(i+1)*binning,:,:], axis=0)\n        elif binning == 1:\n            cds_binned = cds_transposed\n        else:\n            raise Exception(\"Binning must be greater than 0.\")\n        return cds_binned","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:45.795267Z","iopub.execute_input":"2024-08-21T20:28:45.795684Z","iopub.status.idle":"2024-08-21T20:28:45.806249Z","shell.execute_reply.started":"2024-08-21T20:28:45.795646Z","shell.execute_reply":"2024-08-21T20:28:45.804909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%writefile planet_feature_selection.py\nfrom planet_data_airs import PlanetDataAIRSCH0\nimport numpy as np\n\nclass PlanetFeaturesAIRSCH0():\n    def __init__(self,\n                 planet_index: str,\n                 binning: int):\n        self.planet_index = planet_index\n        self.binning = binning\n        \n        self.planet_data = PlanetDataAIRSCH0(self.planet_index)\n        \n        self.planet_data.apply_analogue_to_digital_conversion()\n        self.planet_data.evaluate_cds_of_processed_signal_data()\n        self.cds = self.planet_data.get_cds_of_processed_signal_data_after_binning(binning=self.binning)\n        self.all_light_cds = self.cds.mean(axis=(1,2))  # I.e., sum over space and wavelength\n        \n        self.edge_detection_filter = np.array([-1, 0, 1])\n        self.convolution = np.convolve(self.all_light_cds, self.edge_detection_filter, 'same')\n        self.convolution[0] = np.nan  # First and last datapoints are not valid\n        self.convolution[-1] = np.nan  # First and last datapoints are not valid\n        self.lower_frame_of_transit = np.nanargmax(self.convolution)\n        self.upper_frame_of_transit = np.nanargmin(self.convolution)\n        \n        self.average_during_transit_all = self.all_light_cds[self.lower_frame_of_transit:self.upper_frame_of_transit].mean()\n        self.average_outside_transit_all = np.concatenate((self.all_light_cds[:self.lower_frame_of_transit],self.all_light_cds[self.upper_frame_of_transit:]), axis=0).mean()\n        \n        x = np.concatenate([np.arange(0, self.lower_frame_of_transit), np.arange(self.upper_frame_of_transit, self.all_light_cds.shape[0])])\n        x_grid = np.arange(0, self.all_light_cds.shape[0])\n        y = np.concatenate([self.all_light_cds[:self.lower_frame_of_transit],self.all_light_cds[self.upper_frame_of_transit:]])\n        self.a, self.b, self.c, self.d, self.e = np.polyfit(x, y, 4)\n        self.trend = (self.a * (x_grid ** 4.)) + (self.b * (x_grid ** 3.)) + (self.c * (x_grid ** 2.)) + (self.d * (x_grid ** 1.)) + self.e\n        self.all_light_cds_detrended = self.all_light_cds - self.trend + self.e  # We do not want to remove offset\n    \n    def return_percentage_of_photons_absorbed_summed_over_all_wavelengths(self):\n        self.average_during_transit_all = self.all_light_cds_detrended[self.lower_frame_of_transit:self.upper_frame_of_transit].mean()\n        self.average_outside_transit_all = np.concatenate((self.all_light_cds_detrended[:self.lower_frame_of_transit],self.all_light_cds_detrended[self.upper_frame_of_transit:]), axis=0).mean()\n        return 1. - (self.average_during_transit_all / self.average_outside_transit_all)\n    \n    def return_percentage_of_photons_absorbed_at_wavelength(self):\n        self.average_during_transit = self.cds[self.lower_frame_of_transit:self.upper_frame_of_transit,:,:].mean(axis=(0,2))\n        self.average_outside_transit = np.concatenate((self.cds[:self.lower_frame_of_transit,:,:],self.cds[self.upper_frame_of_transit:,:,:]), axis=0).mean(axis=(0,2))\n        return 1. - self.average_during_transit / self.average_outside_transit","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:45.807736Z","iopub.execute_input":"2024-08-21T20:28:45.808136Z","iopub.status.idle":"2024-08-21T20:28:45.824372Z","shell.execute_reply.started":"2024-08-21T20:28:45.808103Z","shell.execute_reply":"2024-08-21T20:28:45.823039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Please change the planet index to have a go yourself!","metadata":{}},{"cell_type":"code","source":"# Pick one from:\ntrain_labels[\"planet_id\"].head(5)","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:45.825965Z","iopub.execute_input":"2024-08-21T20:28:45.826347Z","iopub.status.idle":"2024-08-21T20:28:45.841958Z","shell.execute_reply.started":"2024-08-21T20:28:45.826307Z","shell.execute_reply":"2024-08-21T20:28:45.840407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from planet_feature_selection import PlanetFeaturesAIRSCH0\n# Let us explore a random planets data\nplanet_id_index = 159506197\nplanet_features = PlanetFeaturesAIRSCH0(planet_id_index, binning=30)","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:45.843678Z","iopub.execute_input":"2024-08-21T20:28:45.844077Z","iopub.status.idle":"2024-08-21T20:28:49.309287Z","shell.execute_reply.started":"2024-08-21T20:28:45.844042Z","shell.execute_reply":"2024-08-21T20:28:49.308206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(3, figsize=(6, 10))\nfig.suptitle(f\"Planet {planet_id_index}:\")\nax[0].set_title(\"All light CDS (Averaged over spatial and wavelength)\")\nax[0].plot(planet_features.all_light_cds, 'k-')\nax[0].set_ylabel(\"Average Count\")\n\nx = np.arange(0, planet_features.all_light_cds.shape[0])\nax[1].set_title(\"Edge detection convolution\")\nax[1].plot(x, planet_features.convolution, 'b--')\nax[1].plot(x[planet_features.lower_frame_of_transit], planet_features.convolution[planet_features.lower_frame_of_transit], 'bs', label=\"Beginning of transit\")\nax[1].plot(x[planet_features.upper_frame_of_transit], planet_features.convolution[planet_features.upper_frame_of_transit], 'rs', label=\"End of transit\")\nax[1].set_ylabel(\"Convolution\")\nax[1].legend()\n\nax[2].set_title(\"All light CDS with trend\")\nax[2].plot(planet_features.all_light_cds, 'k-', label=\"All light cds\")\nax[2].plot(planet_features.trend, 'g--', label='Trend of data outside of transit')\nmin_val = np.min(planet_features.all_light_cds)\nmax_val = np.max(planet_features.all_light_cds)\nax[2].plot([planet_features.lower_frame_of_transit, planet_features.lower_frame_of_transit], [min_val, max_val], 'b--', label=\"Detected beginnning of transit\")\nax[2].plot([planet_features.upper_frame_of_transit, planet_features.upper_frame_of_transit], [min_val, max_val], 'r--', label=\"Detected end of transit\")\nax[2].set_xlabel(\"Binned CDS time step\")\nax[2].set_ylabel(\"Average Count\")\nax[2].legend()\n#plt.savefig(r'/kaggle/working/Feature_Engineering_Method.jpg')","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:49.310561Z","iopub.execute_input":"2024-08-21T20:28:49.310882Z","iopub.status.idle":"2024-08-21T20:28:50.1915Z","shell.execute_reply.started":"2024-08-21T20:28:49.310856Z","shell.execute_reply":"2024-08-21T20:28:50.190302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ratio_absorbed = planet_features.return_percentage_of_photons_absorbed_summed_over_all_wavelengths()\n\nplt.plot(planet_features.all_light_cds_detrended, 'k-')\nplt.title(\"Detrended CDS\\n \\nTotal Light Absorbed (via Feature Engineering): {:.4f}%\".format(ratio_absorbed * 100.))\nx_before_transit = np.arange(0, planet_features.lower_frame_of_transit)\nx_during_transit = np.arange(planet_features.lower_frame_of_transit, planet_features.upper_frame_of_transit)\nx_after_transit = np.arange(planet_features.upper_frame_of_transit, planet_features.all_light_cds_detrended.shape[0])\nplt.plot(planet_features.all_light_cds_detrended, 'k-')\nplt.plot(x_before_transit, np.ones_like(x_before_transit) * planet_features.average_outside_transit_all, 'g--', label=\"Mean outside of transit\")\nplt.plot(x_after_transit, np.ones_like(x_after_transit) * planet_features.average_outside_transit_all, 'g--')\nplt.plot(x_during_transit, np.ones_like(x_during_transit) * planet_features.average_during_transit_all, 'b--', label=\"Mean during of transit\")\nplt.xlabel(\"Binned CDS time step\")\nplt.ylabel(\"Detrended Average Count\")\nplt.legend()\n#plt.savefig(r'/kaggle/working/Feature_Engineering_Result.jpg')","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:50.193028Z","iopub.execute_input":"2024-08-21T20:28:50.193417Z","iopub.status.idle":"2024-08-21T20:28:50.585933Z","shell.execute_reply.started":"2024-08-21T20:28:50.193384Z","shell.execute_reply":"2024-08-21T20:28:50.584616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"The approximate number of photons absorbed is: {:.4f}%\".format(ratio_absorbed * 100.))","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:50.588982Z","iopub.execute_input":"2024-08-21T20:28:50.589387Z","iopub.status.idle":"2024-08-21T20:28:50.59641Z","shell.execute_reply.started":"2024-08-21T20:28:50.589349Z","shell.execute_reply":"2024-08-21T20:28:50.594852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Cleaning AIRS Data","metadata":{}},{"cell_type":"code","source":"PATH_FOLDER = '/kaggle/input/ariel-data-challenge-2024/'\nNUMBER_OF_FRAMES = 11250\nNUMBER_OF_SPATIAL_POINTS = 32\nNUMBER_OF_WAVELENGTH_POINTS = 356\nLOWER_WAVELENGTH_INDEX_CUTOFF = 39\nUPPER_WAVELENGTH_INDEX_CUTOFF = 321\nNUMBER_OF_WAVELENGTH_POINTS_FILTERED = UPPER_WAVELENGTH_INDEX_CUTOFF - LOWER_WAVELENGTH_INDEX_CUTOFF\n\nTRAIN_ADC_INFO = pd.read_csv(os.path.join(PATH_FOLDER, 'train_adc_info.csv')).set_index('planet_id')\nAXIS_INFO = pd.read_parquet('/kaggle/input/ariel-data-challenge-2024/axis_info.parquet')\nDT_AIRS  = AXIS_INFO['AIRS-CH0-integration_time'].dropna().values","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:50.597807Z","iopub.execute_input":"2024-08-21T20:28:50.598243Z","iopub.status.idle":"2024-08-21T20:28:50.62562Z","shell.execute_reply.started":"2024-08-21T20:28:50.598202Z","shell.execute_reply":"2024-08-21T20:28:50.624029Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%writefile get_index.py\n\nimport os\nimport numpy as np\n\n## From GORDONYIP notebook\ndef process_get_index(files):\n    index = []\n    for file in files :\n        file_name = file.split('/')[-1]\n        if file_name.split('_')[0] == 'AIRS-CH0' and file_name.split('_')[1] == 'signal.parquet':\n            file_index = os.path.basename(os.path.dirname(file))\n            index.append(int(file_index))\n    index = np.array(index)\n    index = np.sort(index)\n    \n    return index","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:50.627294Z","iopub.execute_input":"2024-08-21T20:28:50.627798Z","iopub.status.idle":"2024-08-21T20:28:50.636469Z","shell.execute_reply.started":"2024-08-21T20:28:50.627746Z","shell.execute_reply":"2024-08-21T20:28:50.635004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%writefile load_and_clean_airs_data.py\n\nimport os\nimport glob\nimport numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\n\nfrom planet_feature_selection import PlanetFeaturesAIRSCH0\nimport get_index as get_index\n\nPATH_FOLDER = '/kaggle/input/ariel-data-challenge-2024/'\nNUMBER_OF_FRAMES = 11250\nNUMBER_OF_SPATIAL_POINTS = 32\nNUMBER_OF_WAVELENGTH_POINTS = 356\nLOWER_WAVELENGTH_INDEX_CUTOFF = 39\nUPPER_WAVELENGTH_INDEX_CUTOFF = 321\nNUMBER_OF_WAVELENGTH_POINTS_FILTERED = UPPER_WAVELENGTH_INDEX_CUTOFF - LOWER_WAVELENGTH_INDEX_CUTOFF\n\nTRAIN_LABELS_DF = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/train_labels.csv')\n\ndef process_load_and_clean_airs_data(tmp_directory: str,\n                                     path_out: str,\n                                     limit_samples_loaded: int,\n                                     binning: int = 30) -> str:\n    \n    if not os.path.exists(tmp_directory):\n        os.makedirs(tmp_directory)\n        print(f\"Directory {tmp_directory} created.\")\n    else:\n        print(f\"Directory {tmp_directory} already exists.\")\n    \n    if not os.path.exists(path_out):\n        os.makedirs(path_out)\n        print(f\"Directory {path_out} created.\")\n    else:\n        print(f\"Directory {path_out} already exists.\")\n    \n    files = glob.glob(os.path.join(PATH_FOLDER + 'train/', '*/*'))\n    \n    file_load_limit = 4 * limit_samples_loaded  # There are four data files per planet\n    index = get_index.process_get_index(files[:file_load_limit])\n    number_of_frames_after_binning = int(0.5 * NUMBER_OF_FRAMES / binning)\n\n    feature_list = [\"planet_index\", \"total_light_absorbed\", \"w1_label\"]\n    airs_df = pd.DataFrame(0, index=np.arange(limit_samples_loaded), columns=feature_list)  # Define empty DF\n    airs_df[\"planet_index\"] = airs_df[\"planet_index\"].astype('int')\n    airs_df[\"total_light_absorbed\"] = airs_df[\"total_light_absorbed\"].astype('float')\n    airs_df[\"w1_label\"] = airs_df[\"w1_label\"].astype('float')\n\n    for i in tqdm(range(limit_samples_loaded)):\n        planet_features = PlanetFeaturesAIRSCH0(index[i], binning=binning)\n        airs_df.iloc[i, airs_df.columns.get_loc(\"planet_index\")] = index[i]\n        airs_df.iloc[i, airs_df.columns.get_loc(\"total_light_absorbed\")] = planet_features.return_percentage_of_photons_absorbed_summed_over_all_wavelengths()\n        airs_df.iloc[i, airs_df.columns.get_loc(\"w1_label\")] = TRAIN_LABELS_DF[TRAIN_LABELS_DF[\"planet_id\"] == index[i]][\"wl_1\"].item()\n        del planet_features\n\n    airs_df.to_csv(path_out + 'data_train.csv')\n    del airs_df\n    return path_out + 'data_train.csv'\n    ","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:50.638101Z","iopub.execute_input":"2024-08-21T20:28:50.639389Z","iopub.status.idle":"2024-08-21T20:28:50.654879Z","shell.execute_reply.started":"2024-08-21T20:28:50.639344Z","shell.execute_reply":"2024-08-21T20:28:50.653752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import load_and_clean_airs_data as load_and_clean_airs_data\nTMP_DIRECTORY = '/kaggle/tmp/data_light_raw/'\nPATH_OUT = '/kaggle/working/'\nLIMIT_SAMPLES_LOADED = 673  # Leaving last sample as does not work with bin size\nBINNING = 30\ndata_train_path = load_and_clean_airs_data.process_load_and_clean_airs_data(tmp_directory=TMP_DIRECTORY,\n                                                                            path_out=PATH_OUT,\n                                                                            limit_samples_loaded=LIMIT_SAMPLES_LOADED,\n                                                                            binning=BINNING)\ndata_train = pd.read_csv(data_train_path)","metadata":{"execution":{"iopub.status.busy":"2024-08-21T20:28:50.656672Z","iopub.execute_input":"2024-08-21T20:28:50.657123Z","iopub.status.idle":"2024-08-21T21:22:29.459543Z","shell.execute_reply.started":"2024-08-21T20:28:50.657089Z","shell.execute_reply":"2024-08-21T21:22:29.457496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Correlation with Train Labels","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import r2_score\nr2_score_on_data = r2_score(data_train[\"w1_label\"].to_numpy(), data_train[\"total_light_absorbed\"].to_numpy())\nplt.title(\"Extracted Feature for each planet vs w1_label\\n R2 Score: {:.4f}\".format(r2_score_on_data))\nplt.plot(data_train[\"total_light_absorbed\"] * 100., data_train[\"w1_label\"] * 100., 'k.')\nplt.ylabel(r\"w1_label [%]\")\nplt.xlabel(r\"Total Light Absorbed (via Feature Engineering) [%]\")\n#plt.savefig(r'/kaggle/working/Correlation.jpg')","metadata":{"execution":{"iopub.status.busy":"2024-08-21T21:22:29.463301Z","iopub.execute_input":"2024-08-21T21:22:29.463943Z","iopub.status.idle":"2024-08-21T21:22:30.74729Z","shell.execute_reply.started":"2024-08-21T21:22:29.463876Z","shell.execute_reply":"2024-08-21T21:22:30.745711Z"},"trusted":true},"execution_count":null,"outputs":[]}]}