{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<h1 style=\"text-align: center;\"><b>EDA Simplified:<span style=\"color:\n    #87CEEB;\"> Identifying Contrails to Reduce Global Warming\n</span> (English/Mandarin Version)</b></h1>\n\n<h1 style=\"text-align: center;\"><b>简化探索性数据分析：<span style=\"color:\n    #87CEEB;\"> 识别轨迹以减少全球变暖\n</span> (中英文版)</b></h1>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #87CEEB; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Introduction (介绍)</h2>\n\nLet's say that you start walking or biking into a specific place outside like a restaurant or a library on a glorious sunny day with clear skies and some wispy clouds. All of a sudden, you see long vertical lines forming out of nowhere and you are curious about what caused the formation of the midair white streak. Later on, you found out that a specific airplane created a long vertical white line in the middle of the beautiful sky, which is known as a contrail.\n\n假设在晴朗的天空和几缕云彩，你开始步行或骑自行车到外面的特定地方，如餐厅或图书馆。突然间，你看到不知从哪里冒出一条长长的垂直线，你很好奇是什么原因导致了半空中的白色条纹的形成。后来，你发现一架特定的飞机在美丽的天空中间划出了一条垂直的长白线，这就是众所周知的轨迹。","metadata":{}},{"cell_type":"markdown","source":"Contrails are clouds of ice crystals that were generated by the aircraft engine exhaust. It made global warming worse when a specific contrail line trapped heat in the atmosphere. However, researchers created machine learning models for predicting when a specific contrail forms and how much warming it causes though they need to validate their models with satellite imagery from NOAA GOES-16. With some hurdles ahead of solving the contrail problem, Google Research created this competition for machine learners to identify whether a contrail will form so that they'll prevent their formation from each aircraft. And as we head into identifying contrails from satellite imagery, let's soar into our data analysis into detecting contrails in the skies!\n\n凝结尾迹是由飞机发动机排气产生的冰晶云。当一条特定的轨迹线将热量困在大气中时，它使全球变暖变得更糟。然而，研究人员创建了机器学习模型来预测特定尾迹何时形成以及它会导致多少变暖，尽管他们需要使用 NOAA GOES-16 的卫星图像来验证他们的模型。由于在解决轨迹问题之前存在一些障碍，Google Research 为机器学习者举办了这场比赛，以确定是否会形成轨迹，以便他们防止每架飞机形成轨迹。当我们开始从卫星图像中识别航迹时，让我们深入分析我们的数据以检测天空中的航迹！\n\n<center>\n    <img src=\"https://storage.googleapis.com/kaggle-media/competitions/Google-Contrails/waterdroplets.png\" width=500>\n    <figcaption style=\"color: gray;\">Condensation trails form when water vapor condenses on soot and freeze into ice particles. Credit: Imperial College.\n    当水蒸气在烟灰上凝结并冻结成冰粒时，就会形成凝结痕迹。图片来源：帝国理工学院。</figcaption>\n</center>","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #87CEEB; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Imports and Data Setup (导入和数据设置)</h2>\n\nFor visualizing the data in the Contrail Identification competition, we load the pandas module as pd and the numpy module as np for using data science and possibly linear algebra. Then, we import the matplotlib module's pyplot attribute as plt as well as the plotly module's express function as px for plotting out the graphs from specific data. Last but not least, we import the display module from the IPython module followed by the animation module from the matplotlib module for visualizing the animations of each satellite imagery.\n\n为了可视化轨迹识别竞赛中的数据，我们：\n1. 将 `pandas` 模块加载为 `pd`，将 `numpy` 模块加载为 `np`，以便使用数据科学和可能的线性代数。\n2. 将 `matplotlib` 模块的 `pyplot` 属性作为 `plt` 导入，并将 `plotly` 模块的 `express` 函数作为 `px` 导入，以根据特定数据绘制图形。\n3. 从 `IPython` 模块导入`显示`模块，然后从 `matplotlib` 模块导入动画模块，以可视化每个卫星图像的动画。","metadata":{}},{"cell_type":"code","source":"# Data Science and Linear Algebra (数据科学和线性代数)\nimport pandas as pd\nimport numpy as np\n\n# Plotting Graphs (绘制图表)\nimport matplotlib.pyplot as plt\nimport plotly.express as px\n\n# Visualizing Animations (可视化动画)\nfrom IPython import display\nfrom matplotlib import animation","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:27.009373Z","iopub.execute_input":"2023-06-05T15:51:27.010034Z","iopub.status.idle":"2023-06-05T15:51:27.820264Z","shell.execute_reply.started":"2023-06-05T15:51:27.010002Z","shell.execute_reply":"2023-06-05T15:51:27.819415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following importing the necessary modules for our data visualization on contrail detection, we found out that there are json files in this contrail detection competition, the train_metadata, and the valid_metadata. Nevertheless, we create two dataframes, train_df, and valid_df, to read out the json file with the pd module's read_json file, setting the file directory that leads to the train_metadata and valid_metadata json files. Thenceforth, we use the head function in the two dataframes for displaying the first five rows.\n\n导入轨迹检测数据可视化的必要模块后，我们发现在这个轨迹检测比赛中有json文件，`train_metadata`和`valid_metadata`。尽管如此，我们还是创建了两个数据框，`train_df` 和 `valid_df`，用 `pd` 模块的 `read_json` 文件读出 json 文件，设置通向 `train_metadata` 和 `validation_metadata` json 文件的文件目录。此后，我们在两个数据框中使用 `head` 函数来显示前五行。","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_json(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/train_metadata.json\")\nvalid_df = pd.read_json(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/validation_metadata.json\")","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:27.821625Z","iopub.execute_input":"2023-06-05T15:51:27.822634Z","iopub.status.idle":"2023-06-05T15:51:28.207227Z","shell.execute_reply.started":"2023-06-05T15:51:27.822602Z","shell.execute_reply":"2023-06-05T15:51:28.206291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.210048Z","iopub.execute_input":"2023-06-05T15:51:28.210505Z","iopub.status.idle":"2023-06-05T15:51:28.24087Z","shell.execute_reply.started":"2023-06-05T15:51:28.21045Z","shell.execute_reply":"2023-06-05T15:51:28.239685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid_df.head()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.243402Z","iopub.execute_input":"2023-06-05T15:51:28.244126Z","iopub.status.idle":"2023-06-05T15:51:28.261111Z","shell.execute_reply.started":"2023-06-05T15:51:28.244084Z","shell.execute_reply":"2023-06-05T15:51:28.259904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h3 style=\"background-color: #87CEEB; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 8px; border-radius: 15px;\">Basic Analysis of the Dataframes (数据框的基本分析)</h3>\n\nNow that we built our train_df and valid_df dataframes, let's start our basic analysis of the two dataframes! First, let's search the number of overall data entities in the train_df and valid_df dataframes with the len function.\n\n现在我们构建了 `train_df` 和 `valid_df` 数据帧，让我们开始对这两个数据帧的基本分析！首先，让我们使用 `len` 函数搜索 `train_df` 和 `valid_df` 数据帧中的整体数据实体数。","metadata":{}},{"cell_type":"code","source":"len(train_df)","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.262936Z","iopub.execute_input":"2023-06-05T15:51:28.263361Z","iopub.status.idle":"2023-06-05T15:51:28.271901Z","shell.execute_reply.started":"2023-06-05T15:51:28.263325Z","shell.execute_reply":"2023-06-05T15:51:28.270678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(valid_df)","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.273668Z","iopub.execute_input":"2023-06-05T15:51:28.274079Z","iopub.status.idle":"2023-06-05T15:51:28.281786Z","shell.execute_reply.started":"2023-06-05T15:51:28.274044Z","shell.execute_reply":"2023-06-05T15:51:28.280622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we saw from finding the number of data in each of the two dataframes, we tallied 20529 entities in the train_df dataframe, while we calculated 1856 entities in the valid_df dataframe. Specifically, the large number of data in the train_df dataframe hinted to us that there are a lot of satellite imagery for the machine learning models to detect the formation of contours while the other data in the valid_df dataframe were used for validating the data for detecting contours.\n\n从我们从两个数据帧中的数据数量中看到的情况来看，我们在 `train_df` 数据帧中统计了 **20529** 个实体，而在 `valid_df` 数据帧中计算了 **1856** 个实体。具体来说，`train_df` 数据框中的大量数据暗示我们有大量卫星图像供机器学习模型检测轮廓的形成，而 `valid_df` 数据框中的其他数据用于验证检测数据轮廓。","metadata":{}},{"cell_type":"markdown","source":"Let's now find the number of missing values in both dataframes! We simply use the isna function to search for missing values in a specific dataframe and then we use the sum function to calculate the overall missing values by the dataframe's column.\n\n现在让我们找出两个数据框中缺失值的数量！我们简单地使用 `isna` 函数来搜索特定数据框中的缺失值，然后我们使用 `sum` 函数来计算数据框列的整体缺失值。","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.283593Z","iopub.execute_input":"2023-06-05T15:51:28.284014Z","iopub.status.idle":"2023-06-05T15:51:28.303394Z","shell.execute_reply.started":"2023-06-05T15:51:28.283977Z","shell.execute_reply":"2023-06-05T15:51:28.302067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.30529Z","iopub.execute_input":"2023-06-05T15:51:28.305744Z","iopub.status.idle":"2023-06-05T15:51:28.31756Z","shell.execute_reply.started":"2023-06-05T15:51:28.305707Z","shell.execute_reply":"2023-06-05T15:51:28.316253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Fantastic! There are no missing values shown in all of the columns from the train_df and the valid_df dataframes. That means all of the data was gathered from the satellite imagery by researchers since they wanted to validate their models for identifying contrails forming in the air.\n\n极好的！ `train_df` 和 `valid_df` 数据帧的所有列中都没有显示缺失值。这意味着所有数据都是研究人员从卫星图像中收集的，因为他们想验证他们的模型以识别空气中形成的尾迹。","metadata":{}},{"cell_type":"markdown","source":"Last but not least, let's visualize the number of columns in the train_df and valid_df dataframes! All we have to do is to apply the shape attribute to the two dataframes specified and then grab the last index that is listed as 1.\n\n最后但同样重要的是，让我们可视化 `train_df` 和 `valid_df` 数据帧中的列数！我们所要做的就是将形状属性应用于指定的两个数据帧，然后获取列为 **1** 的最后一个索引。","metadata":{}},{"cell_type":"code","source":"train_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.319355Z","iopub.execute_input":"2023-06-05T15:51:28.319778Z","iopub.status.idle":"2023-06-05T15:51:28.328566Z","shell.execute_reply.started":"2023-06-05T15:51:28.319739Z","shell.execute_reply":"2023-06-05T15:51:28.327411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"valid_df.shape[1]","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.33355Z","iopub.execute_input":"2023-06-05T15:51:28.334036Z","iopub.status.idle":"2023-06-05T15:51:28.341863Z","shell.execute_reply.started":"2023-06-05T15:51:28.333997Z","shell.execute_reply":"2023-06-05T15:51:28.340736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, there are 7 columns in both train_df and valid_df dataframes. Additionally, the same number of columns we noticed in both dataframes implied to us that the data is separated for the machine learning models to train and then validate their accuracy and confidence in detecting contrails.\n\n正如我们所见，`train_df` 和 `valid_df` 数据帧中都有 **7** 列。此外，我们在两个数据框中注意到的相同数量的列暗示我们，数据是分开的，以便机器学习模型进行训练，然后验证它们在检测尾迹方面的准确性和信心。","metadata":{}},{"cell_type":"markdown","source":"As we complete our small basic analysis on the train_df and valid_df dataframes, let's begin plotting the data from both dataframes as we push forward into our data analysis in the contrail detection competition!\n\n当我们完成对 `train_df` 和 `valid_df` 数据帧的小型基本分析时，让我们开始绘制来自两个数据帧的数据，因为我们将在轨迹检测竞赛中推进我们的数据分析！ ","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #87CEEB; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Visualizing the train_df Dataframe (可视化 \"train_df\" 数据框)</h2>\n\nWhen we reached into visualizing the train_df dataframe, we came across the training metadata for each record of satellite imagery, as it held the timestamps and the projection parameters to recreate the satellite images. Aside from our brief and short summary of the train_df dataframe's overview, let's explain the columns one by one!\n\n当我们开始可视化 `train_df` 数据帧时，我们遇到了每条卫星图像记录的训练元数据，因为它包含时间戳和投影参数以重新创建卫星图像。除了我们对 `train_df` 数据框概述的简短总结之外，让我们一一解释这些列！\n\n* **record_id**: Identifier for each record of satellite imagery. (每条卫星图像记录的标识符。)\n* **projection_wkt**: The well-known text of each projection. (每个投影的众所周知的文本。)\n* **row_min**, **col_min**: The minimum of the row and column from each projection. (每个投影的行和列的最小值。)\n* **col_min**, **col_size**: The size of each row and column from each projection. (每个投影的每一行和每一列的大小。)\n* **timestamp**: The date of when the projection was created. (创建投影的日期。)","metadata":{}},{"cell_type":"markdown","source":"Now that we completed explaining the seven columns in the train_df dataframe with concise detail, let's distribute the data from the record_id column into the histogram! To get started, we characterize the fig variable to plot a histogram plot with the px module's histogram function, setting the train_df dataframe as the data for the histogram plot, followed by the x parameter to the record_id column for configuring the x-axes, and the marginal parameter to \"box\" for plotting the box-plot on the top of our histogram graph. After we created the histogram from the fig variable figure, we apply the show function into the fig variable figure for displaying the graph to the code cell output below.\n\n现在我们已经简明扼要地详细解释了 `train_df` 数据框中的七列，让我们将 `record_id` 列中的数据分配到直方图中！首先，我们描述 `fig` 变量以使用 `px` 模块的直方图函数绘制直方图，将 `train_df` 数据帧设置为直方图的数据，然后将 `x` 参数设置为 `record_id` 列以配置 \"x\" 轴，以及“box”的边际参数，用于在我们的直方图顶部绘制箱线图。从 `fig` 变量数字创建直方图后，我们将 `show` 函数应用于 `fig` 变量数字以将图形显示到下面的代码单元格输出中。","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(train_df, x=\"record_id\", marginal=\"box\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:28.343639Z","iopub.execute_input":"2023-06-05T15:51:28.344082Z","iopub.status.idle":"2023-06-05T15:51:30.367216Z","shell.execute_reply.started":"2023-06-05T15:51:28.344046Z","shell.execute_reply":"2023-06-05T15:51:30.366034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After we compiled the histogram based on the data distribution of the record_id column, we noted that the bins in the histogram graph displayed a uniform distribution, as most of the bins had the same counts. Additionally, the highest data counted is in the range between 6.6x10<sup>18</sup> and 6.8x10<sup>18</sup>, with 483 entities, while the other range between 9.2x10<sup>18</sup> and 9.4x10<sup>18</sup> is the least counted data, with 61 entities. To be specific, the uniform distribution we saw in the record_id column histogram gave us clues that most of the records in the contrails competition data contained an average of around 390 to 480 satellite images, while the last records contained a few satellite images possibly because of unfinished recordings.\n\n在我们根据 `record_id` 列的数据分布编译直方图后，我们注意到直方图中的 \"bin\" 显示均匀分布，因为大多数 \"bin\" 具有相同的计数。此外，计数最高的数据在 **6.6x10**<sup><b>18</b></sup> 和 **6.8x10**<sup><b>18</b></sup> 之间的范围内，有 **483** 个实体，而 **9.2x10**<sup><b>18</b></sup> 和 **9.4x10**<sup><b>18</b></sup> 之间的另一个范围是计数最少的数据，有 **61** 个实体。具体来说，我们在 `record_id` 列直方图中看到的均匀分布为我们提供了轨迹比赛数据中的大部分记录平均包含大约 **390** 到 **480** 张卫星图像的线索，而最后的记录包含一些卫星图像可能是因为未完成的录音。","metadata":{}},{"cell_type":"markdown","source":"Let's now visualize and distribute the row_min and row_size columns separately into two histograms in one subplot! Before we begin plotting this out, we import two additional modules, the plotly module's graph_objects attribute as go as well as the make_subplots function from the plotly module's subplots attribute for creating the subplots with Plotly. \n\n现在让我们将 `row_min` 和 `row_size` 列可视化并分别分布到一个子图中的两个直方图中！在开始绘制之前，我们导入了两个额外的模块，`plotly` 模块的 `graph_objects` 属性以及 `plotly` 模块的 `subplots` 属性中的 `make_subplots` 函数，用于使用 \"Plotly\" 创建子图。\n\nAfter we import the two necessary modules, we characterize the fig variable to create our subplot graph with the make_subplots function, setting the rows parameter to 1 and the col parameter to 2 for configuring our subplot with a row and two columns. We then apply the histogram graphs with the add_trace function into the fig variable figure twice but individually, setting the go module's Histogram function that has the x parameter set to the train_df dataframe's standalone row_min and row_size columns and the name parameter to the name of the column that is specified from the x parameter for configuring two of our histograms with the labeled traces followed by arranging the row parameter to 1 for placing both histograms in the single row of our subplot, and the col parameter to 1 and 2 separately for placing each histogram plot into two columns of the subplot. Finally, we use the show function in the fig variable figure for displaying the graph below the code cell.\n\n导入两个必要的模块后，我们使用 `make_subplots` 函数对 `fig` 变量进行特征化以创建我们的子图，将 `rows` 参数设置为 **1** 并将 `col` 参数设置为 **2** 以使用一行和两列配置我们的子图。然后，我们将带有 `add_trace` 函数的直方图图形应用到 `fig` 变量 `figure` 中两次，但是是单独的，将 `go` 模块的 `Histogram` 函数将 `x` 参数设置为 `train_df` 数据帧的独立 `row_min` 和 `row_size` 列，并将 `name` 参数设置为列的名称这是从 `x` 参数指定的，用于配置我们的两个带有标记轨迹的直方图，然后将 row 参数设置为 **1** 以将两个直方图放置在子图的单行中，并将 `col` 参数分别设置为 **1** 和 **2** 以放置每个直方图绘制成子图的两列。最后，我们使用 `fig` 变量 `figure` 中的 `show` 函数在代码单元格下方显示图形。","metadata":{}},{"cell_type":"code","source":"import plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\n\nfig = make_subplots(rows=1, cols=2)\n\nfig.add_trace(\n    go.Histogram(x=train_df[\"row_min\"], name=\"row_min\"),\n    row=1, col=1\n)\n\nfig.add_trace(\n    go.Histogram(x=train_df[\"row_size\"], name=\"row_size\"),\n    row=1, col=2\n)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.36838Z","iopub.execute_input":"2023-06-05T15:51:30.368769Z","iopub.status.idle":"2023-06-05T15:51:30.419395Z","shell.execute_reply.started":"2023-06-05T15:51:30.368731Z","shell.execute_reply":"2023-06-05T15:51:30.418405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the row_min histogram, we found out that the tall ascending bars from the train_df dataframe's row_min column, were separate from each other along with the adjacent bars ascending from the near right of the diagram, as the separate tall ascending bars from the left to the right of the histogram merely exhibited a left-skew distribution. Aside from observing the looks of the row_min column, the highest data range counted is from 4.6 million to 4.7 million, with 1632 entities, while the range from 2.8 million to 2.9 million is counted the least with only one entity. In addition, the right skew distribution we saw from the separate ascending bars implied to us that there are most specific high minimum values in some projections of the satellite image records.\n\n从 `row_min` 直方图中，我们发现来自 `train_df` 数据帧的 `row_min` 列的高上升条与从图的右侧附近上升的相邻条彼此分开，因为单独的高上升条从左到右直方图的右侧仅呈现左偏分布。抛开`row_min`列的样子，统计的数据范围最高的是**460**万到**470**万，有**1632**个实体，而**280**万到**290**万的范围统计最少，只有一个实体。此外，我们从单独的上升条中看到的右偏分布向我们暗示，在卫星图像记录的某些投影中存在最具体的高最小值。\n\nMeanwhile from the row_size data distribution on the right of the subplot graph, we noticed that on the left of the histogram chart, there are jagged bars ascending constantly, as it clearly showed a left-skewed data distribution. Besides, the highest data range counted is from -1980 to -1978, with 665 entities, while the ranges from -2146 to -2144 and -1864 and -1862 had only one entity, making them the least-counted data ranges. Overall, the left-skewed distribution we glimpsed from the row_size column gave us clues that most of the row size values are moderately-big for displaying a specific projection, despite having negative values.\n\n同时从子图右侧的`row_size`数据分布中，我们注意到在左侧的柱状图上，有不断上升的锯齿状柱状图，明显呈现出左偏的数据分布。此外，计数最高的数据范围是 **-1980** 到 **-1978**，有**665**个实体，而 **-2146** 到 **-2144** 和 **-1864** 和 **-1862**只有一个实体，是计数最少的数据范围。总体而言，我们从 `row_size` 列中瞥见的左偏分布为我们提供了线索，即尽管具有负值，但大多数行大小值对于显示特定投影来说都是适度大的。","metadata":{}},{"cell_type":"markdown","source":"Now let's find the relations between the row_min and row_size columns into a 2D heatmap histogram graph! We again define the fig variable figure into creating our heatmap histogram chart with the density_heatmap function from the px module, setting the train_df dataframe as the data for the 2D histogram, followed by configuring the x and y parameters to the row_min and row_size columns for specifying the x and y axes of our graph. Once completed, we display our graph in the output cell by using the show function to the fig variable figure.\n\n现在让我们将 `row_min` 和 `row_size` 列之间的关系找到一个 \"2D\" 热图直方图！我们再次将 `fig` 变量 \"figure\" 定义为使用 `px` 模块中的 `density_heatmap` 函数创建我们的热图直方图图表，将 `train_df` 数据帧设置为 \"2D\" 直方图的数据，然后将 `x` 和 `y` 参数配置到 `row_min` 和 `row_size` 列以指定我们图表的 \"x\" 和 \"y\" 轴。完成后，我们通过对 `fig` 变量 \"figure\" 使用 `show` 函数在输出单元格中显示我们的图形。","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(train_df, x=\"row_min\", y=\"row_size\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.420568Z","iopub.execute_input":"2023-06-05T15:51:30.421596Z","iopub.status.idle":"2023-06-05T15:51:30.524572Z","shell.execute_reply.started":"2023-06-05T15:51:30.421557Z","shell.execute_reply":"2023-06-05T15:51:30.523628Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Following from configuring our 2D density histogram graph based on the row_min and row_size columns in the train_df dataframe, we spotted on how the left-skews from the row_min and row_size columns were clearly mapped, as it formed a curved right-triangle shape on the right of the heatmap diagram. Overall, the row_min column's range from 500K to 1M and row_size's column range between -1960 and -1950 was counted the most, with 566 data entities, while there are some ranges in the row_size and row_min columns was counted once, making them have the least-counted data. In conclusion, the visible left-skewed data we found from the right of the diagram above gave us clues that most projections have high minimum values in their rows while having negatively moderately-big values in their sizes.\n\n在基于 `train_df` 数据帧中的 `row_min` 和 `row_size` 列配置我们的二维密度直方图之后，我们发现了 `row_min` 和 `row_size` 列的左偏如何被清楚地映射，因为它在右侧形成了一个弯曲的直角三角形的热图。总体来说，`row_min`列的范围从**500K**到**1M**和`row_size`的列范围**-1960**到**-1950**被统计最多，有**566**个数据实体，而`row_size`和`row_min`列有一些范围被统计一次，最少- 统计数据。总之，我们从上图右侧发现的可见左偏数据为我们提供了线索，即大多数投影在其行中具有较高的最小值，同时在其大小中具有负中等大的值。","metadata":{}},{"cell_type":"markdown","source":"Besides probing the row_min and row_size columns, let's proceed towards distributing and then analyzing the col_min and col_size columns into another two histograms inside a subplot! First of all, we create and assign the fig variable graph to build our subplot graph with the make_subplots function, setting the rows and cols parameters to 1 and 2 for configuring a row and two columns into our subplot. Next, we use the add_trace function twice and individually into the fig variable graph to add our two histogram graphs, containing the go module's Histogram function that has the x parameter configured to the train_df dataframe's standalone col_min and col_size columns as well as the name parameter to the name of the column specified in the x parameter, the row parameter to 1, and the col parameter to 1 and 2 separately for placing the histogram charts to two columns in the subplot chart. Last but not least, we apply the show function to the fig variable chart for displaying the graph underneath the code cell.\n\n除了探测 `row_min` 和 `row_size` 列之外，让我们继续将 `col_min` 和 `col_size` 列分布并分析到子图中的另外两个直方图中！首先，我们创建并分配 `fig` 变量图以使用 `make_subplots` 函数构建我们的子图，将 `rows` 和 `cols` 参数设置为 **1** 和 **2** 以将一行和两列配置到我们的子图中。接下来，我们两次使用 `add_trace` 函数并分别将其添加到 `fig` 变量图中以添加我们的两个直方图，其中包含 `go` 模块的 `Histogram` 函数，其 `x` 参数配置为 `train_df` 数据帧的独立 `col_min` 和 `col_size` 列以及 `name` 参数`x`参数中指定的列名，`row`参数为**1**，`col`参数分别为**1**和**2**，用于将直方图图表放置到\"subplot\"图表中的两列。最后但同样重要的是，我们将 `show` 函数应用于 `fig` 变量图表，以在代码单元格下方显示图形。","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=2)\n\nfig.add_trace(go.Histogram(\n    x=train_df[\"col_min\"],\n    name=\"col_min\"\n), row=1, col=1)\n\nfig.add_trace(go.Histogram(\n    x=train_df[\"col_size\"],\n    name=\"col_size\"\n), row=1, col=2)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.525816Z","iopub.execute_input":"2023-06-05T15:51:30.526828Z","iopub.status.idle":"2023-06-05T15:51:30.56892Z","shell.execute_reply.started":"2023-06-05T15:51:30.526788Z","shell.execute_reply":"2023-06-05T15:51:30.567828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the data distribution of the col_min column on the left of our subplot above, we found out that there are jagged bars in some places of the histogram, as it nearly showed the symmetrical distribution. Aside from the looks of each graph, the highest data range counted is from 560k to 570k, with around 699 entities, while the other data field from between 810k and 820k with 19 entities. Overall, the jagged bins we saw in the col_size data distribution hinted to us that there are numeric variations of each column size in each projection.\n\n从上面子图左侧`col_min`列的数据分布，我们发现直方图的某些地方有锯齿状的条形，几乎呈对称分布。除了每张图的外观，统计的最高数据范围是从 **560k** 到 **570k**，有大约 **699** 个实体，而其他数据字段在 **810k** 到 **820k** 之间，有 **19** 个实体。总的来说，我们在 `col_size` 数据分布中看到的锯齿状 \"bin\" 暗示我们每个投影中每个列大小的数值变化。\n\nMeanwhile for the col_size data distribution in another histogram on the right of our subplot, we noticed that the uneven bars ascend rapidly in the near-left of the histogram, as it clearly exhibits a left-skew distribution. In other way of explaination, the highest data count is in the range from 1954 to 1956, with 824 entities, while there are other ranges that had one entity, making them as the least-counted data. Furthermore, the left-skew distribution we noticed from the right histogram graph in the subplot hinted to us that there are some moderately-big size in some columns in each projection.\n\n同时，对于子图右侧另一个直方图中的 `col_size` 数据分布，我们注意到直方图左侧附近的不均匀条快速上升，因为它明显呈现左偏分布。换句话说，**1954**年到**1956**年的数据数量最多，有**824**个实体，而其他范围只有一个实体，是最少的数据。此外，我们从子图中的右侧直方图中注意到的左偏分布向我们暗示，在每个投影中的某些列中存在一些中等大小的尺寸。","metadata":{}},{"cell_type":"markdown","source":"Since we visualized the col_min and col_size columns individually in a subplot, let's proceed with visualizing the relations between the two columns into another 2D histogram heatmap plot! We simply build our fig variable figure and then define it to create a density heatmap graph with the density_heatmap function from the px module, setting the train_df dataframe as the data for the chart, followed by the x and y parameters to col_min and col_size for specifying the x and y axes. Last but not least, we apply the show function to the fig variable figure for exhibiting the graph at the bottom of the code cell.\n\n由于我们在子图中分别可视化了 `col_min` 和 `col_size` 列，让我们继续将两列之间的关系可视化到另一个 \"2D\" 直方图热图图中！我们简单地构建我们的 `fig` 变量图，然后定义它以使用 `px` 模块中的 `density_heatmap` 函数创建密度热图图，将 `train_df` 数据帧设置为图表的数据，然后将 `x` 和 `y` 参数设置为 `col_min` 和 `col_size` 以指定`x` 和 `y` 轴。最后但同样重要的是，我们将 `show` 函数应用于 `fig` 变量 \"figure\" 以在代码单元格底部显示图形。","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(train_df, x=\"col_min\", y=\"col_size\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.570323Z","iopub.execute_input":"2023-06-05T15:51:30.571223Z","iopub.status.idle":"2023-06-05T15:51:30.644361Z","shell.execute_reply.started":"2023-06-05T15:51:30.571182Z","shell.execute_reply":"2023-06-05T15:51:30.643329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see from the relationship between the col_min and the col_size columns in the heatmap graph, we glimpsed how the heated areas from the left of the diagram expand to the right of the diagram, as we noticed the magenta areas in this place. In other words, the col_min range from 200k to 150k and the col_size range from 1950 to 1960 is the most common relation we noticed, with 959 data entities, while there are some ranges in the col_size and col_min columns that had one data entity, making them least common. Overall, the expansion of the heated areas we saw from the 2D histogram between the col_min and col_size columns gave us some clues that there are some projections in which their minimum columns vary numerically while having moderately-big column sizes.\n\n从热图中 `col_min` 和 `col_size` 列之间的关系可以看出，我们瞥见了图表左侧的加热区域如何扩展到图表的右侧，因为我们注意到这个地方的洋红色区域。换句话说，`col_min` 范围从 **200k** 到 **150k** 和 `col_size` 范围从 **1950** 到 **1960** 是我们注意到的最常见的关系，有 **959** 个数据实体，而 `col_size` 和 col_min 列中有一些范围只有一个数据实体，使得他们最不常见。总体而言，我们从 `col_min` 和 `col_size` 列之间的 \"2D\" 直方图中看到的加热区域的扩展给了我们一些线索，表明在某些投影中，它们的最小列在数值上有所不同，同时具有适度大的列大小。","metadata":{}},{"cell_type":"markdown","source":"As we finished analyzing the col_min and col_size columns' relationship above, we completed on visualizing most of the data from the train_df dataframe! What's next ahead is that we are going to study the data from the valid_df dataframe, as the data in there was different from the train_df dataframe's data. \n\n当我们分析完上面 `col_min` 和 `col_size` 列的关系时，我们完成了 `train_df` 数据帧中大部分数据的可视化！接下来我们将研究来自 `valid_df` 数据帧的数据，因为其中的数据与 `train_df` 数据帧的数据不同。","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #87CEEB; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Visualizing the valid_df Dataframe (可视化 \"valid_df\" 数据框)</h2>\n\nSince we visualized the overall data from the train_df dataframe, we proceed to probe another dataframe that is used for validation when a specific machine learning model finished training after each epoch, which is the valid_df dataframe. Surprisingly, the valid_df dataframe had the same number of columns as the train_df dataframe meaning that we don't have to repeat the detailing of each column in this dataframe. Without further ado, let's take off to investigate the valid_df dataframe!\n\n由于我们可视化了来自 `train_df` 数据帧的整体数据，因此我们继续探索另一个数据帧，该数据帧用于在特定机器学习模型在每个时期后完成训练时进行验证，即 `valid_df` 数据帧。令人惊讶的是，`valid_df` 数据帧的列数与 `train_df` 数据帧的列数相同，这意味着我们不必重复此数据帧中每一列的详细信息。事不宜迟，让我们开始研究 `valid_df` 数据帧！","metadata":{}},{"cell_type":"markdown","source":"First, let's distribute and then analyze the record_id column into a standard histogram graph without a box plot graph on top of it! To start this process, we characterize the fig variable figure into creating a histogram with the px module's histogram function, inputting the valid_df dataframe as the data for the histogram chart, followed by configuring the x parameter to the record_id column for arranging the x-axes in the chart. Last but not least, we use the show function in the fig variable figure for displaying the graph beneath the code cell.\n\n首先，让我们将 `record_id` 列分布并分析为标准直方图，上面没有箱线图！为了开始这个过程，我们将 `fig` 变量 \"figure\" 表征为使用 `px` 模块的 `histogram` 函数创建直方图，输入 `valid_df` 数据帧作为直方图图表的数据，然后将 `x` 参数配置到 `record_id` 列以排列 \"x\" 轴在图表中。最后但同样重要的是，我们使用 `fig` 变量 \"figure\" 中的 `show` 函数在代码单元格下方显示图形。","metadata":{}},{"cell_type":"code","source":"fig = px.histogram(valid_df, x=\"record_id\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.645765Z","iopub.execute_input":"2023-06-05T15:51:30.646157Z","iopub.status.idle":"2023-06-05T15:51:30.71503Z","shell.execute_reply.started":"2023-06-05T15:51:30.646125Z","shell.execute_reply":"2023-06-05T15:51:30.71384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"How strange. As we can see from the data distribution of the record_id column, we noticed that the bins from the histogram chart were jagged while descending a bit, nearly indicating this as a uniform distribution. By way of explanation, the most counted data in the record_id column is in the range from 2x10<sup>18</sup> to 2.5x<sup>18</sup>, with 118 data entities, while the least counted data is in the other range in between 9x10<sup>18</sup> and 9.5x10<sup>18</sup>, with 48 data entities. Additionally, the uniform distribution we mostly noticed from the data distribution of the record_id column hinted to us that almost all of the recording projections for validation contained an average of between 83 to 118 satellite images, while the last ranges contained a few satellite images because of possibly unfinished recordings.\n\n多么奇怪。正如我们从 `record_id` 列的数据分布中看到的那样，我们注意到直方图中的 \"bin\" 在下降时呈锯齿状，几乎表明这是一种均匀分布。解释一下，`record_id`列中统计最多的数据在**2x10**<sup><b>18</b></sup>到**2.5x**<sup><b>18</b></sup>之间，有**118**个数据实体，统计最少的是另一个范围在 **9x10**<sup><b>18</b></sup> 和 **9.5x10**<sup><b>18</b></sup> 之间，有 **48** 个数据实体。此外，我们主要从 `record_id` 列的数据分布中注意到的均匀分布向我们暗示，几乎所有用于验证的记录投影平均包含 **83** 到 **118** 个卫星图像，而最后一个范围包含一些卫星图像，因为可能是未完成的录音。","metadata":{}},{"cell_type":"markdown","source":"Let's now move on to distributing the row_min and row_size columns into two separate histograms in a subplot! Once again, we create our fig variable figure and then use it to create our subplot graph with the make_subplots function, setting the rows parameter to 1 and the cols parameter to 2 for making our subplot have one row and two columns. Thenceforth, we apply the add_trace function into the fig variable figure twice but separately, setting the go module's Histogram function that has the x parameter set to the valid_df dataframe's row_min and row_size columns individually as well as the name parameter to the name of the column specified from the x parameter, followed by configuring the row parameter to 1 and the col parameter to 1 and 2 one by one for placing the histogram traces to a specific position in the subplot graph. Lastly, we display our subplot graph below the code cell by applying the show function to the fig variable figure.\n\n现在让我们继续将 `row_min` 和 `row_size` 列分布到子图中的两个单独的直方图中！再一次，我们创建了我们的 `fig` 变量图形，然后用它来创建我们的子图，使用 `make_subplots` 函数，将 `rows` 参数设置为 **1**，将 `cols` 参数设置为 **2** 以使我们的子图具有一行和两列。此后，我们将 `add_trace` 函数应用于 `fig` 变量 `figure` 两次，但分别进行，将 `go` 模块的 `Histogram` 函数设置为 `x` 参数分别设置为 `valid_df` 数据帧的 `row_min` 和 `row_size` 列，并将 `name` 参数设置为指定列的名称从 `x` 参数开始，然后将 `row` 参数配置为 **1**，将 `col` 参数配置为 **1** 和 **2**，用于将直方图轨迹放置到 `subplot` 图中的特定位置。最后，我们通过将 `show` 函数应用于 `fig` 变量 `figure` 来在代码单元下方显示我们的子图。","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=2)\n\nfig.add_trace(go.Histogram(x=valid_df['row_min'], name='row_min'), row=1, col=1)\nfig.add_trace(go.Histogram(x=valid_df['row_size'], name='row_size'), row=1, col=2)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.717265Z","iopub.execute_input":"2023-06-05T15:51:30.718499Z","iopub.status.idle":"2023-06-05T15:51:30.796311Z","shell.execute_reply.started":"2023-06-05T15:51:30.718458Z","shell.execute_reply":"2023-06-05T15:51:30.795249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the row_min data column we graphed in the left of the subplot, we glimpsed that there's a data peak on the far right of the diagram, as we noticed that the data bins ascend abruptly in the near middle of the chart, making this as a left-skewed data distribution. Aside from the looks of the row_min data distribution, the highest-counted data is in the range from 3.5M to 3.9M, with 201 entities, while the other range spanning between -4M and -3.6M is the lowest-counted data, with 20 data entities. Additionally, the left-skewed distribution we observed from the row_min chart gave us some clues that most of the projections' rows in the validation metadata have more height that have positive values than the ones that have negative values.\n\n从我们在子图左侧绘制的 `row_min` 数据列中，我们瞥见图表的最右侧有一个数据峰值，因为我们注意到数据箱在图表的中间附近突然上升，这使得它成为左偏数据分布。除了 `row_min` 数据分布的外观外，计数最高的数据在 **3.5M** 到 **3.9M** 范围内，有 **201** 个实体，而另一个介于 **-4M** 和 **-3.6M** 之间的数据是计数最低的数据，有**20** 个数据实体。此外，我们从 `row_min` 图表中观察到的左偏分布为我们提供了一些线索，即验证元数据中的大多数投影行具有更多具有正值的高度而不是具有负值的行。\n\nMeanwhile from what we noticed in the row_size distribution, we found out that there was a tall, jagged data peak on the near-left of the diagram, as it clearly showed the left-skewed data distribution. Specifically, the highest data recorded in the row_size column is between -1980 and -1975, with 143 entities, while the range from -1860 to -1855 has the least-counted data, with around 3 entities. Furthermore, the left-skewed data distribution we noticed from the row_size distribution gave us hints that there is some moderately-small size in some columns in each projection recorded from the valid_df dataframe despite having negative values.\n\n同时，从我们在 `row_size` 分布中注意到的情况，我们发现在图表的左侧附近有一个高大的锯齿状数据峰，因为它清楚地显示了左偏数据分布。具体来说，`row_size`列中记录的数据最多的是 **-1980** 到 **-1975**之间，有**143**个实体，而 **-1860** 到 **-1855** 之间的数据最少，有**3**个左右的实体。此外，我们从 row_size 分布中注意到的左偏数据分布给了我们提示，尽管具有负值，但在从 `valid_df` 数据帧记录的每个投影中的某些列中存在一些适度较小的尺寸。","metadata":{}},{"cell_type":"markdown","source":"Now that we visualized the row_size and row_min columns individually in two histograms, let's visualize the relation between the two columns in a 2D heatmap histogram chart! We simply create the fig variable figure and then define it into making the 2D histogram heatmap chart with the density_heatmap function provided by the px module, setting the valid_df dataframe as the data input for the graph, followed by characterizing the x and y parameters to the row_size and row_min columns for configuring the x and y axes. With that completed, we use the show function to the fig variable figure setup as we're displaying the chart in the code output beneath the cell.\n\n现在我们在两个直方图中分别可视化了 `row_size` 和 `row_min` 列，让我们在 \"2D\" 热图直方图中可视化两列之间的关系！我们简单的创建`fig`变量`figure`，然后定义成使用`px`模块提供的`density_heatmap`函数制作\"2D histogram heatmap\"图表，设置`valid_df` \"dataframe\"作为图形的数据输入，然后将`x`和`y`参数表征为`row_size` 和 `row_min` 列用于配置 \"x\" 和 \"y\" 轴。完成后，当我们在单元格下方的代码输出中显示图表时，我们将 `show` 函数用于 `fig` 变量图形设置。","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(valid_df, x=\"row_size\", y=\"row_min\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.797937Z","iopub.execute_input":"2023-06-05T15:51:30.798572Z","iopub.status.idle":"2023-06-05T15:51:30.875587Z","shell.execute_reply.started":"2023-06-05T15:51:30.798533Z","shell.execute_reply":"2023-06-05T15:51:30.874489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From what we observed in the 2D histogram heatmap chart about the relationship between the row_size and the row_min columns, we noticed that there are some orange and magenta areas that nearly formed a diagonal line that is pointing down. Additionally, the highest data relation is in the range between 0 to 900k in the row_min column and -1960 to -1940.01 in the row_size column, with 140 entities, while the ranges from -4M to -3.1M in the row_min column and from -1980 to -1960.01 in the row_size column is the lowest data relation, with 4 entities. In summary, the downward diagonal line we noticed from the relations between the row_size and row_min columns gave us suggestions that some data rows have more height with positive values in the row_min column and have negatively moderate-small values in the row_size column.\n\n从我们在 \"2D\" 直方图热图中观察到的 `row_size` 和 `row_min` 列之间的关系，我们注意到有一些橙色和洋红色区域几乎形成了一条向下的对角线。此外，最高数据关系在 row_min 列中的 **0** 到 **900k** 和 `row_size` 列中的 **-1960** 到 **-1940.01** 之间的范围内，有 **140** 个实体，而 `row_min` 列中的范围从 **-4M** 到 **-3.1M** 和 - `row_size` 列中的 **1980** 到 **-1960.01** 是最低的数据关系，有 **4** 个实体。总而言之，我们从 `row_size` 和 `row_min` 列之间的关系中注意到的向下对角线给了我们一些建议，即某些数据行在 `row_min` 列中具有更高的正值，而在 `row_size` 列中具有负中小值。","metadata":{}},{"cell_type":"markdown","source":"Let's shift our gears into visualizing the data from the col_min and col_size into two individual histograms in one subplot! First, we characterize the fig variable figure by creating our subplot chart with the make_subplots function, setting the rows parameter to 1 and the cols parameter to 2. Next, we add our charts into our subplot chart with the add_trace function plugged to the fig variable figure, placing the go module's Histogram function that has the x parameter configured to the valid_df dataframe's col_min and col_size columns separately and the name parameter to the column name arranged by the x parameter, followed by setting the row parameter to 1 and the col parameter to 1 and 2 individually for placing our graphs into two different columns of the subplot graph within the same row. Finally, we exhibit our graph underneath the code cell by using the show function to the fig variable figure.\n\n让我们换档，将来自 `col_min` 和 `col_size` 的数据可视化为一个子图中的两个单独的直方图！首先，我们通过使用 `make_subplots` 函数创建我们的子图来表征 `fig` 变量图，将 `rows` 参数设置为 **1**，将 `cols` 参数设置为 **2**。接下来，我们将我们的图表添加到我们的子图图表中，并将 `add_trace` 函数插入 `fig` 变量如图，将`go`模块的`Histogram`函数配置了`x`参数分别放在`valid_df` \"dataframe\"的`col_min`和`col_size`列，`name`参数设置为`x`参数排列的列名，然后设置`row`参数为**1**，`col`参数为**1** 和 **2** 分别用于将我们的图形放入同一行内子图的两个不同列中。最后，我们通过对 `fig` 变量 \"figure\" 使用 `show` 函数在代码单元下方显示我们的图表。","metadata":{}},{"cell_type":"code","source":"fig = make_subplots(rows=1, cols=2)\n\nfig.add_trace(go.Histogram(x=valid_df[\"col_min\"], name=\"col_min\"), row=1, col=1)\nfig.add_trace(go.Histogram(x=valid_df[\"col_size\"], name=\"col_size\"), row=1, col=2)\n\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.877214Z","iopub.execute_input":"2023-06-05T15:51:30.877628Z","iopub.status.idle":"2023-06-05T15:51:30.909638Z","shell.execute_reply.started":"2023-06-05T15:51:30.877584Z","shell.execute_reply":"2023-06-05T15:51:30.908428Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the col_min column, we glimpsed how there's a jagged data peak that is mostly in the middle of the diagram, as it almost showed a symmetrical distribution. Aside from the looks of the col_min data distribution, the data range that was counted the most is between 560k and 579.99k, with 117 entities, while the other range spanning from 640k to 659.99k was counted the least, with 6 entities. Additionally, the jagged, symmetrical distribution we saw from the col_min column implied to us that there are moderate values in the minimum size of some columns in some projections.\n\n从 `col_min` 列中，我们瞥见了一个锯齿状的数据峰值，该峰值主要位于图表的中间，因为它几乎呈对称分布。除了 `col_min` 数据分布的外观外，计数最多的数据范围在 **560k** 到 **579.99k** 之间，有 **117** 个实体，而另一个范围从 **640k** 到 **659.99k** 计数最少，有 **6** 个实体。此外，我们从 `col_min` 列中看到的锯齿状对称分布向我们暗示，在某些投影中，某些列的最小大小具有适中的值。\n\nOn the other graph to the right of the subplot, which showed the visualization of the col_size column, we found out that the bars on the far-left of the graph mostly ascended constantly until the near-middle of the chart, displaying the left-skewed distribution of the col_size column. Besides, the most common data of the col_size column is in the range spanning from 1950 to 1954.99, with 204 entities, while the other range between 2055 to 2059.99 has the least common data, with only one entity. Funnily enough, the right-skewed data distribution we noticed from the col_size column gave us the overall summary that there are fewer values in some projection's column sizes.\n\n在子图右侧的另一个图表上，显示了 `col_size` 列的可视化，我们发现图表最左侧的条形大部分不断上升，直到图表的中部附近，显示左侧 - `col_size` 列的偏态分布。此外，`col_size` 列最常见的数据在 **1950** 到 **1954.99** 之间的范围内，有 **204** 个实体，而 **2055** 到 **2059.99** 之间的另一个范围内的数据最不常见，只有一个实体。有趣的是，我们从 `col_size` 列中注意到的右偏数据分布为我们提供了总体总结，即某些投影的列大小中的值较少。","metadata":{}},{"cell_type":"markdown","source":"Once we completed visualizing the col_size and col_min columns in two histograms, let's now probe the relation between the two data columns into the 2D histogram heatmap graph! To do that, we create the fig variable figure and assign it to the px module's density_heatmap function, setting the valid_df dataframe as the data for the graph, followed by arranging the x and y parameters to the col_min and col_size columns for specifying the x and y axes in the graph. Lastly, we apply the show function to the fig variable figure so that we'll display the graph in the output below the code cell.\n\n一旦我们完成了两个直方图中 `col_size` 和 `col_min` 列的可视化，现在让我们将两个数据列之间的关系探查到 \"2D\" 直方图热图中！为此，我们创建 `fig` 变量 \"figure\" 并将其分配给 `px` 模块的 `density_heatmap` 函数，将 `valid_df` 数据帧设置为图形的数据，然后将 `x` 和 `y` 参数安排到 `col_min` 和 `col_size` 列以指定 \"x\" 和图中的 \"y\" 轴。最后，我们将 `show` 函数应用于 `fig` 变量 \"figure\" 以便我们将在代码单元格下方的输出中显示图形。","metadata":{}},{"cell_type":"code","source":"fig = px.density_heatmap(valid_df, x=\"col_min\", y=\"col_size\")\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.911217Z","iopub.execute_input":"2023-06-05T15:51:30.911631Z","iopub.status.idle":"2023-06-05T15:51:30.971454Z","shell.execute_reply.started":"2023-06-05T15:51:30.911592Z","shell.execute_reply":"2023-06-05T15:51:30.970427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After running the above code cell on creating the col_min and col_size heatmap graph, we noticed that there's a curved right triangle shape at the bottom of the diagram as it clearly showed a left skew distribution from the col_size column. To be specific, the highest data area is in the ranges between 200k and 299.9k in the col_min column and 1940 to 1959.9 in the col_size column with 192 entities, while the other range from 700k to 799.9k in the col_min column and from 2040 to 2059.9 in the col_size column is the lowest data area, with only one entity. Additionally, the curved right triangle we glimpsed from the relationship between the col_min and the col_size columns gave us cues that there are some moderate values in the minimum of the column size and some small values in the overall column size from some projections.\n\n在运行上面的代码单元创建 `col_min` 和 `col_size` 热图后，我们注意到图表底部有一个弯曲的直角三角形，因为它清楚地显示了 `col_size` 列的左偏分布。具体来说，最高的数据区域在`col_min`列中的**200k**到**299.9k**之间和`col_size`列中的**1940**到**1959.9**之间，有**192**个实体，而其他范围在`col_min`列中的**700k**到**799.9k**和**2040**到**1959.9**之间`col_size` 列中的 **2059.9** 是最低的数据区，只有一个实体。此外，我们从 `col_min` 和 `col_size` 列之间的关系中瞥见的弯曲直角三角形给了我们一些线索，即在某些投影中，列大小的最小值有一些中等值，而总列大小有一些小值。","metadata":{}},{"cell_type":"markdown","source":"After visualizing the col_min and col_size relations in the heatmap 2D histogram graph, we finally completed our visualization of the valid_df dataframe! Now that we visualized most of the data from the train_df and the valid_df dataframes, let's take off into visualizing the contrails in the sky!\n\n在热图二维直方图中可视化 `col_min` 和 `col_size` 关系后，我们终于完成了 `valid_df` 数据帧的可视化！现在我们已经可视化了来自 `train_df` 和 `valid_df` 数据帧的大部分数据，让我们开始可视化天空中的轨迹！","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #87CEEB; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Probing the Contrails in the Blue (探索蓝色的凝结尾迹)</h2>\n\nPreviously in our EDA journey through detecting contrails in the sky, we visualized the train_df and valid_df dataframes, which it is the training and validation metadata of the records in each satellite imagery. Right now, we now arrive at the final section of our data visualization on contrail detection, which is visualizing the contrails from npy masks and image files. Additionally, the satellite images were originally recorded by the GOES-16 ABI as it is publicly available in Google Cloud Storage. And since contrails are too easy to spot with temporal context, there's an image sequence at 10-minute intervals.\n\n之前在我们通过检测天空轨迹的 \"EDA\" 旅程中，我们可视化了 `train_df` 和 `valid_df` 数据帧，它们是每个卫星图像中记录的训练和验证元数据。现在，我们到达了关于航迹检测的数据可视化的最后一部分，即可视化来自 \"npy\" 掩码和图像文件的航迹。此外，卫星图像最初由 \"GOES-16 ABI\" 记录，因为它在谷歌云存储中公开可用。由于轨迹很容易在时间背景下被发现，因此每隔 **10** 分钟就有一个图像序列。","metadata":{}},{"cell_type":"markdown","source":"To kickstart into this section, we open the directory which indicates the record id to any three bands specified in the npy file and read it in binary mode with \"rb\" as the f variable, then we define the variable with the band name to load the f variable with the np module's load function. We then open and then read it in binary mode with \"rb\" to the human pixel and human individual masks in the npy file as the f variable individually, and thenceforth characterize the human_pixel_mask and human_individual_mask variables separately to load the f variable file with the np module's load function.\n\n为了快速进入本节，我们打开 \"npy\" 文件中指定的任何三个波段的指示记录 \"ID\" 的目录，并以“rb”作为 `f` 变量以二进制模式读取它，然后我们定义要加载波段名称的变量带有 ``np`` 模块加载函数的 `f` 变量。。然后我们打开然后用“rb”二进制方式读取到\"npy\"文件中的\"human pixel\"和\"human individual masks\"作为f变量分别作为f变量，然后分别对`human_pixel_mask`和`human_individual_mask`变量进行特征化，将np加载到f变量文件中模块的加载函数。","metadata":{}},{"cell_type":"code","source":"# The bands (乐队)\nwith open(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/train/1023779081293051446/band_15.npy\", \"rb\") as f:\n    band_15 = np.load(f)\nwith open(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/train/1023779081293051446/band_14.npy\", \"rb\") as f:\n    band_14 = np.load(f)\nwith open(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/train/1023779081293051446/band_11.npy\", \"rb\") as f:\n    band_11 = np.load(f)\n    \n# The human individual and pixel masks (人类个体和像素掩码)\nwith open(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/train/1023779081293051446/human_pixel_masks.npy\", \"rb\") as f:\n    human_pixel_mask = np.load(f)\nwith open(\"/kaggle/input/google-research-identify-contrails-reduce-global-warming/train/1023779081293051446/human_individual_masks.npy\", \"rb\") as f:\n    human_individual_mask = np.load(f)","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:30.973294Z","iopub.execute_input":"2023-06-05T15:51:30.973673Z","iopub.status.idle":"2023-06-05T15:51:31.299776Z","shell.execute_reply.started":"2023-06-05T15:51:30.973642Z","shell.execute_reply":"2023-06-05T15:51:31.298801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's combine the bands into a false-color image! For viewing the contrails in the GOES satellite imagery, it used the \"ash\" color scheme so that it views the thin cirrus, including contrails, as it appears in the image as dark blue. Aside from the overview of the false color image, we characterize the three variables into three separate tuples: the t11_bounds variable to a tuple with the values 243 to 303, the cloud_top_tdiff_bounds variable to the values -4 and 5 in a tuple, and the tdiff_bounds variable to the values -4 and 2 in another tuple.\n\n现在让我们将波段组合成假彩色图像！为了查看 \"GOES\" 卫星图像中的尾迹，它使用了“灰”配色方案，以便它可以看到薄卷云，包括尾迹，因为它在图像中显示为深蓝色。除了伪彩色图像的概述之外，我们将这三个变量表征为三个独立的元组：`t11_bounds` 变量为一个值为 **243** 到 **303** 的元组，`cloud_top_tdiff_bounds` 变量为一个元组中的值 **-4** 和 **5**，以及`tdiff_bounds` 变量到另一个元组中的值 **-4** 和 **2**。","metadata":{}},{"cell_type":"code","source":"t11_bounds = (243, 303)\ncloud_top_tdiff_bounds = (-4, 5)\ntdiff_bounds = (-4, 2)","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:31.301009Z","iopub.execute_input":"2023-06-05T15:51:31.301813Z","iopub.status.idle":"2023-06-05T15:51:31.306804Z","shell.execute_reply.started":"2023-06-05T15:51:31.301772Z","shell.execute_reply":"2023-06-05T15:51:31.305577Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We then create our function to map out the data to the range specified as 0 and 1 in an array which we call normalize_range, defining the data and bounds parameters, and then return the division of the difference between the data parameter and the bounds parameter's first slice index as 0 and the difference between the bounds parameter's slice index of 1 and 0. Thenceforth, we characterize the r, g, and b variables to map out the data to the range of 0 and 1 with the normalize_range function we created, individually placing the subtraction of the band_15, band_14 variables, and band_14, band_11 variables as well as the band_14 variable for the data, followed by applying the tdiff_bounds, cloud_top_tdiff_bounds, and t11_bounds variable as the bounds for mapping the data. Lastly, we create the false_color variable to clip the array that has the r, g, and b variables joined by the np module's stack function along with the values 0 and 1 as the min and max to clip to with the clip function from the np module.\n\n然后我们创建我们的函数，将数据映射到我们称为 `normalize_range` 的数组中指定为 **0** 和 **1** 的范围，定义 `data` 和 `bounds` 参数，然后返回数据参数和 `bounds` 参数之间的差异除法切片索引为 **0**，`bounds` 参数的切片索引为 **1** 和 **0** 之间的差值。此后，我们将 `r`、`g` 和 `b` 变量表征为使用我们创建的 `normalize_range` 函数分别将数据映射到 **0** 和 **1** 的范围将 `band_15`、`band_14` 变量和 `band_14`、`band_11` 变量以及数据的 `band_14` 变量相减，然后应用 `tdiff_bounds`、`cloud_top_tdiff_bounds` 和 `t11_bounds` 变量作为映射数据的边界。最后，我们创建 `false_color` 变量来裁剪数组，该数组具有由 `np` 模块的堆栈函数连接的 `r`、`g` 和 `b` 变量以及值 **0** 和 **1** 作为要使用 `np` 的裁剪函数裁剪到的最小值和最大值模块。","metadata":{}},{"cell_type":"code","source":"def normalize_range(data, bounds):\n    return (data - bounds[0]) / (bounds[1] - bounds[0])\n\nr = normalize_range(band_15 - band_14, tdiff_bounds)\ng = normalize_range(band_14 - band_11, cloud_top_tdiff_bounds)\nb = normalize_range(band_14, t11_bounds)\nfalse_color = np.clip(np.stack([r,g,b], axis=2), 0, 1)","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:31.308766Z","iopub.execute_input":"2023-06-05T15:51:31.309455Z","iopub.status.idle":"2023-06-05T15:51:31.344063Z","shell.execute_reply.started":"2023-06-05T15:51:31.309426Z","shell.execute_reply":"2023-06-05T15:51:31.342964Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now here's the fun part: let's soar into visualizing the data of the contrails in the sky! To get started, we characterize the n_times_before variable to 4 for specifying the number of times before projecting the satellite imagery and the img variable to the false_color's slice indexes of the ellipsis for pointing the two numbers and the n_times_before variable for extracting the index. \n\n现在是有趣的部分：让我们开始可视化天空中轨迹的数据！首先，我们将 `n_times_before` 变量表征为 **4**，用于指定投影卫星图像之前的次数，将 `img` 变量表征为省略号的 `false_color` 切片索引，用于指向两个数字，将 `n_times_before` 变量表征为提取索引。\n\nNext, we create our matplotlib chart by using the figure function from the plt module, setting the figsize parameter to 18 by 6 under the parentheses for specifying the width and height of the graph. We then characterize the ax variable to create three subplots with the plt module's subplot function, setting the values 1 and 3 for creating the subplot's one row and three columns followed by individually placing the values 1, 2, and 3 for placing the graphs in each subplot's index. Thenceforth, we display the image into the ax variable with the imshow function three times, setting the img variable for the first one, the pixel_mask variable with the interpolation parameter set to none (just to show no interpolations) for the second one, and the human_pixel_mask variable with the cmap parameter to 'Reds' (colormap), the alpha parameter to .4 (specifies the alpha), and the interpolation parameter to none (no interpolation) for the third one. Lastly, we apply the title for the subplot graph with the set_title function plugged into the ax variable thrice, setting the strings 'False Color Img' for the first title, 'Ground Truth Contrail Mask' for the second one, and 'Contrail Mask on False Color Img' for the third one.\n\n接下来，我们使用 `plt` 模块中的 `figure` 函数创建 \"matplotlib\" 图表，将圆括号下的 `figsize` 参数设置为 **18 x 6** 以指定图形的宽度和高度。然后，我们对 `ax` 变量进行特征化，以使用 `plt` 模块的子图函数创建三个子图，设置值 **1** 和 **3** 以创建子图的一行三列，然后分别设置值 **1**、**2** 和 **3** 以在每个子图中放置图形子图的索引。此后，我们使用 `imshow` 函数将图像显示到 `ax` 变量中三次，第一次设置 `img` 变量，第二次设置 `pixel_mask` 变量并将插值参数设置为 \"none\"（只是为了不显示插值），最后`human_pixel_mask` 变量，其中 `cmap` 参数为“Reds”（颜色图），`alpha` 参数为 **.4**（指定 \"alpha\"），第三个插值参数为 \"none\"（无插值）。最后，我们将带有 `set_title` 函数的子图标题应用到 `ax` 变量三次，第一个标题设置字符串“False Color Img”，第二个标题设置“Ground Truth Contrail Mask”，“Contrail Mask on”第三个的假彩色图像'。","metadata":{}},{"cell_type":"code","source":"n_times_before = 4\nimg = false_color[..., n_times_before]\n\nplt.figure(figsize=(18,6))\nax = plt.subplot(1, 3, 1)\nax.imshow(img)\nax.set_title(\"False Color Img\")\n\nax = plt.subplot(1, 3, 2)\nax.imshow(human_pixel_mask, interpolation='none')\nax.set_title('Ground Truth Contrail Mask')\n\nax = plt.subplot(1, 3, 3)\nax.imshow(img) # <- The third one is like the same as the first one (第三个和第一个一样)\nax.imshow(human_pixel_mask, cmap=\"Reds\", alpha=0.4, interpolation='none')\nax.set_title(\"Contrail Mask on False Color Img\")","metadata":{"execution":{"iopub.status.busy":"2023-06-05T15:51:31.345725Z","iopub.execute_input":"2023-06-05T15:51:31.34608Z","iopub.status.idle":"2023-06-05T15:51:32.41899Z","shell.execute_reply.started":"2023-06-05T15:51:31.346053Z","shell.execute_reply":"2023-06-05T15:51:32.417914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the False Color Image, we espied that there are clouds around the sky while seeing faint thin lines of contrails on the middle bottom of the image. Meanwhile from the Ground Truth Contrail Mask outline, we saw three contrail lines outlined by the ground truth mask on the middle-bottom of the picture. Lastly, for the Contrail Mask on the False Color Image on the right of the subplot, we notice how the three contrail lines were outlined on the false-color image by the ground truth contrail mask. Specifically, the false-color image, the ground truth contrail mask, and the contrail mask on the false color image stored in the subplot graph gave us hints that the satellite imagery projections vary, as some projections and contrail masks show more contrail lines. In contrast, other projections and contrail masks show little or no contrail lines.\n\n从伪彩色图像中，我们看到天空周围有云，同时在图像的中间底部看到微弱的细线轨迹。同时从\"Ground Truth Contrail Mask\"轮廓中，我们看到图片中下部由\"Ground Truth Mask\"勾勒出的三条轨迹线。最后，对于子图右侧的假彩色图像上的轨迹遮罩，我们注意到地面真实轨迹遮罩如何在假彩色图像上勾勒出三条轨迹线。具体来说，伪彩色图像、地面真值航迹掩码和存储在子图中的伪彩色图像上的航迹掩码给了我们卫星图像投影变化的暗示，因为一些投影和航迹掩码显示了更多的轨迹线。相比之下，其他投影和轨迹掩码显示很少或没有轨迹线。","metadata":{}},{"cell_type":"markdown","source":"To conclude our last section of visualizing the contrails in the sky, let's create our animation graph on each satellite projection! First, we characterize the fig variable to create our graph figure with the plt module's figure function, setting the figsize parameter to 6 by 6 in parentheses for setting the width and height of the graph and the im variable to show the image with the imshow function specified by the plt module, setting the false_color variable's slice index of the two values specified as the ellipsis and the value 0 for loading the image.\n\n为了结束可视化天空轨迹的最后一部分，让我们在每个卫星投影上创建动画图！首先，我们使用 `plt` 模块的 `figure` 函数表征 `fig` 变量来创建我们的图形，将括号中的 `figsize` 参数设置为 **6 x 6** 以设置图形的宽度和高度，以及 `im` 变量以使用 `imshow` 函数显示图像由`plt`模块指定，设置`false_color`变量的切片索引为省略号指定的两个值和0值加载图像。\n\nNext, we generate a function to draw out the animation of each satellite projection which we call \"draw\" as we set the i variable as the parameter of the function. Inside of the draw function, we use the set_array function to the im variable for setting the values of the array in the im variable, containing the false_color variable's slice index of the ellipsis (specifies the two values) and the i parameter, thus returning the im variable in square brackets.\n\n接下来，我们生成一个函数来绘制每个卫星投影的动画，我们称之为“draw”，因为我们将 `i` 变量设置为函数的参数。在 `draw` 函数内部，我们使用 `set_array` 函数对 im 变量设置数组在 `im` 变量中的值，包含 `false_color` 变量的省略号切片索引（指定两个值）和 `i` 参数，从而返回方括号中的 `im` 变量。\n\nOutside from characterizing the draw function, we define the anim variable to create our animated chart with the FuncAnimation function from the animation module, placing the fig variable figure for specifying the graph of where to put the animation in, the draw function for drawing the animation of each projection, followed by configuring the frames parameter to the false_color variable's shape's slice index of -1 for specifying the number of frames, the interval parameter to 500 for specifying each pause of each frame, and the blit parameter to True for using blitting when drawing an image. Finally, we close the plot with the plt module's close function and then display our animation with the HTML function from the display function, setting the anim variable being converted to enable the interactive playback with JavaScript with the to_jshtml function inside.\n\n在描述 `draw` 函数之外，我们定义了 `anim` 变量以使用动画模块中的 `FuncAnimation` 函数创建我们的动画图表，放置 `fig` 变量 `figure` 用于指定放置动画的图形，`draw` 函数用于绘制动画每个投影的 `frames` 参数配置为 `false_color` 变量的 `shape` 切片索引 **-1** 用于指定帧数，`interval` 参数配置为 **500** 用于指定每帧的每次暂停，`blit` 参数配置为 `True` 用于在绘制图像。最后，我们使用 `plt` 模块的 `close` 函数关闭绘图，然后使用 `display` 函数中的 `HTML` 函数显示我们的动画，设置正在转换的 `anim` 变量，以启用带有 `to_jshtml` 函数的 \"JavaScript\" 交互式播放。","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize=(6,6))\nim = plt.imshow(false_color[..., 0])\n\ndef draw(i):\n    im.set_array(false_color[..., i])\n    return [im]\n\nanim = animation.FuncAnimation(\n    fig, draw, frames=false_color.shape[-1], interval=500, blit=True\n)\n\nplt.close()\ndisplay.HTML(anim.to_jshtml())","metadata":{"execution":{"iopub.status.busy":"2023-06-05T16:12:12.6474Z","iopub.execute_input":"2023-06-05T16:12:12.647805Z","iopub.status.idle":"2023-06-05T16:12:15.191139Z","shell.execute_reply.started":"2023-06-05T16:12:12.647773Z","shell.execute_reply":"2023-06-05T16:12:15.189922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the first few seconds, we saw the long faint streak of a contrail line in the middle-bottom part of the graph. Then, we glimpsed on seeing two diagonal contrail lines forming below the long faint streak of a contrail line, as all three contrail streaks starts to drift to the bottom right. Specifically, one of the animations we played from a specific contrail projection made us think that there are some cases in which contrails start to form unexpectedly while there are other cases in which a lot of contrail lines bundle up.\n\n从最初的几秒钟开始，我们就在图表的中下部看到了一条长而微弱的尾迹线。然后，我们瞥见两条对角线轨迹形成在一条轨迹线的长而微弱的条纹下方，因为所有三个轨迹轨迹都开始漂移到右下角。具体来说，我们从一个特定的轨迹投影播放的动画之一让我们认为在某些情况下轨迹会意外开始形成，而在其他情况下会有很多轨迹线捆绑在一起。","metadata":{}},{"cell_type":"markdown","source":"And with our short and quick visualization of the contrail lines being done, we are officially done visualizing the data in the contrail detection competition! \n\n通过我们对轨迹线的简短快速可视化，我们正式完成了轨迹检测竞赛中数据的可视化！","metadata":{}},{"cell_type":"markdown","source":"<h2 style=\"background-color: #87CEEB; color: white; padding-right: 100vw; background-size:cover; text-align: center; padding: 10px; border-radius: 15px;\">Conclusion (结论)</h2>\n\nSo what do we learn from visualizing the data in the contrail detection competition? From our data analysis in the train_df and the valid_df dataframes, we learned about the metadata of each satellite projection, as the data were categorized into training and validation. And while visualizing the contrails from each satellite imagery, we acknowledged how contrails form unexpectedly around the sky as well as understanding on where we notice the contrails' location as there are some that were bundled together while others were few and separate from each other. And once we detect contrails accurately in the sky, we'll help the airliners avoid emitting contrails in the sky and sustain the earth's climate before it changes, so that we'll not see glaciers melting followed by more intense storms around the world. \n\n那么我们从轨迹检测竞赛中的数据可视化中学到了什么？通过对 `train_df` 和 `valid_df` 数据帧的数据分析，我们了解了每个卫星投影的元数据，因为数据被分类为训练和验证。在从每张卫星图像中看到凝结尾迹时，我们认识到凝结尾迹是如何在天空中意外形成的，并了解我们注意到凝结尾迹的位置，因为有些凝迹聚集在一起，而另一些则很少且彼此分开。一旦我们准确地检测到天空中的尾迹，我们将帮助客机避免在天空中发射尾迹，并在地球气候发生变化之前维持地球气候，这样我们就不会看到冰川融化后世界各地出现更强烈的风暴。","metadata":{}}]}