{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":101849,"databundleVersionId":13093295,"sourceType":"competition"},{"sourceId":12602965,"sourceType":"datasetVersion","datasetId":7960517}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"***Stratification and Data Exploration for Ariel Data Challenge***\n\nThis notebook documents the investigation of stratification strategies for cross-validation folds in the Ariel data challenge. It shows how we explored the star and planet feature space to decide whether stratified folds would be beneficial. Ultimately, we conclude that simple random folds suffice given the data characteristics.","metadata":{}},{"cell_type":"markdown","source":"***Creating K-Fold Splits with Random Shuffling***\n\nWe start by creating 5 random K-fold splits of our planet samples using sklearn's KFold class. Each fold will serve as a validation set once, providing a robust estimate of model performance. We shuffle the data for randomness but do not stratify by any feature yet, as we first want to investigate if stratification is necessary.\n\nThe generated fold assignments are saved for use throughout the notebook.","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.decomposition import PCA\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import KFold\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.cluster import KMeans\nimport seaborn as sns\nimport warnings\n\n# Suppress FutureWarning related to use_inf_as_na deprecation in seaborn/pandas\nwarnings.filterwarnings(\"ignore\", category=FutureWarning, message=\".*use_inf_as_na.*\")\n\n# Suppress the specific UserWarning about figure layout change in seaborn\nwarnings.filterwarnings(\"ignore\", category=UserWarning, message=\".*figure layout has changed to tight.*\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:34:47.365837Z","iopub.execute_input":"2025-09-27T16:34:47.36618Z","iopub.status.idle":"2025-09-27T16:34:47.373993Z","shell.execute_reply.started":"2025-09-27T16:34:47.366154Z","shell.execute_reply":"2025-09-27T16:34:47.372635Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ## Stratified K-Fold Cross-Validation Setup\n\n# Define data path and dataset\nDATA_PATH = '/kaggle/input/ariel-data-challenge-2025'\nDATASET = 'train'\n\n# Load the planet_ids from star info CSV, converting to int index\nplanet_ids = pd.read_csv(f'{DATA_PATH}/{DATASET}_star_info.csv', index_col='planet_id').index.astype(int)\n\n# Create DataFrame of planet_ids for fold assignment\ndf = pd.DataFrame({'planet_id': planet_ids})\n\n# Initialize KFold with 5 splits, shuffling enabled for randomness and reproducibility\nkf = KFold(n_splits=5, shuffle=True, random_state=42)\n\n# Assign fold numbers to planet_ids by iterating over each fold's validation indices\nfolds = []\nfor fold_index, (_, val_idx) in enumerate(kf.split(planet_ids)):\n    for i in val_idx:\n        folds.append({'planet_id': planet_ids[i], 'fold': fold_index})\n\nfolds_df = pd.DataFrame(folds)\n\n# Save fold assignments to CSV for easy reuse\nfolds_df.to_csv('planet_kfolds.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:25:24.32641Z","iopub.execute_input":"2025-09-27T16:25:24.32684Z","iopub.status.idle":"2025-09-27T16:25:24.352038Z","shell.execute_reply.started":"2025-09-27T16:25:24.326811Z","shell.execute_reply":"2025-09-27T16:25:24.350625Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"***Visualizing Fold Assignments Using PCA***\n\nNext, we visualize how the folds are distributed in the feature space of star data using Principal Component Analysis (PCA). PCA reduces many features down to two principal components for convenient plotting.\n\nPlotting fold membership in this reduced space helps us check if folds contain a good spread of planet types or if some folds cluster unusually. Ideally, folds should overlap well, indicating no major bias introduced by the splitting.","metadata":{}},{"cell_type":"code","source":"# ## PCA Visualization of Fold Assignments\n\n# Load star_info features for all planets (including planets in our folds)\nstar_info = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/train_star_info.csv')\n\n# Load fold assignments created earlier\nfolds = pd.read_csv('/kaggle/working/planet_kfolds.csv')\n\n# Merge star_info and folds on planet_id to align features with fold labels\ndf = pd.merge(star_info, folds, on='planet_id')\n\n# Extract feature matrix and fold labels\nfeatures = df.drop(columns=['planet_id', 'fold']).values\nfold_labels = df['fold'].values\nplanet_ids = df['planet_id'].values","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:25:28.986646Z","iopub.execute_input":"2025-09-27T16:25:28.986976Z","iopub.status.idle":"2025-09-27T16:25:29.015178Z","shell.execute_reply.started":"2025-09-27T16:25:28.986949Z","shell.execute_reply":"2025-09-27T16:25:29.014259Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Reduce dimensionality to 2 principal components for visualization\npca = PCA(n_components=2)\ncomponents = pca.fit_transform(features)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:25:37.53496Z","iopub.execute_input":"2025-09-27T16:25:37.535275Z","iopub.status.idle":"2025-09-27T16:25:37.607932Z","shell.execute_reply.started":"2025-09-27T16:25:37.53525Z","shell.execute_reply":"2025-09-27T16:25:37.606731Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot each fold with distinct color to visualize distribution\nplt.figure(figsize=(8, 6))\ncolors = plt.get_cmap('Set1')\nfor fold in range(5):\n    idx = (fold_labels == fold)\n    plt.scatter(\n        components[idx, 0], components[idx, 1],\n        color=colors(fold), label=f'Fold {fold}', alpha=0.8, edgecolor='none', s=40\n    )\nplt.xlabel(f'PC1 ({pca.explained_variance_ratio_[0]*100:.1f}%)')\nplt.ylabel(f'PC2 ({pca.explained_variance_ratio_[1]*100:.1f}%)')\nplt.legend(title='Fold')\nplt.title('PCA of Star Info Colored by Fold Assignment')\nplt.grid(True)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:25:49.191247Z","iopub.execute_input":"2025-09-27T16:25:49.191606Z","iopub.status.idle":"2025-09-27T16:25:49.666082Z","shell.execute_reply.started":"2025-09-27T16:25:49.191576Z","shell.execute_reply":"2025-09-27T16:25:49.665073Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The graph visualizes the distribution of the 5 cross-validation folds across the first two principal components (PC1 and PC2) of the star feature space. Each fold is represented by a distinct color.\n\nWe observe that the folds overlap reasonably well in this 2D projection, indicating that the data points in each fold span similar regions of the feature space. This good mixing suggests that random KFold splitting has created balanced and diverse folds without severe clustering or segregation.\n\nSuch overlap supports the choice to use random folds instead of more complex stratification, as it helps ensure that each fold is representative of the overall data distribution.","metadata":{}},{"cell_type":"markdown","source":"***Exploratory Data Analysis of Star and Planet Features***\n\nTo understand potential sources of variance or stratification criteria, we analyze the key astrophysical parameters of stars and planets. We begin by standardizing features and exploring their pairwise relationships through plots.\n\nApplying PCA and KMeans clustering reveals if sub-groups or clusters naturally exist in the data.\n\nThis investigation also informs future model development and data processing.","metadata":{}},{"cell_type":"code","source":"# ## Exploratory Data Analysis of Star Features\n\n# Load starinfo dataset containing relevant astrophysical parameters\nstarinfo = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/train_star_info.csv')\n\n# Selected features for analysis\nfeatures = ['Rs', 'Ms', 'Ts', 'Mp', 'P', 'sma', 'i']\n\n# Standardize features to zero mean and unit variance for clustering and PCA\nscaler = StandardScaler()\nscaled_features = scaler.fit_transform(starinfo[features])","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:30:28.710562Z","iopub.execute_input":"2025-09-27T16:30:28.711261Z","iopub.status.idle":"2025-09-27T16:30:28.729119Z","shell.execute_reply.started":"2025-09-27T16:30:28.711221Z","shell.execute_reply":"2025-09-27T16:30:28.727991Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Pairplot shows feature distributions and pairwise relationships with KDE diagonals\nsns.pairplot(starinfo[features], diag_kind='kde', plot_kws={'alpha':0.5})\nplt.suptitle(\"Pairwise Distributions of Starinfo Parameters\", y=1.02)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:34:53.458094Z","iopub.execute_input":"2025-09-27T16:34:53.458467Z","iopub.status.idle":"2025-09-27T16:35:06.590625Z","shell.execute_reply.started":"2025-09-27T16:34:53.458442Z","shell.execute_reply":"2025-09-27T16:35:06.589261Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply PCA and print explained variance ratio for dimensionality reduction\npca = PCA(n_components=2)\npca_result = pca.fit_transform(scaled_features)\nprint(f\"PCA Explained Variance Ratio: {pca.explained_variance_ratio_}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:33:02.479911Z","iopub.execute_input":"2025-09-27T16:33:02.480196Z","iopub.status.idle":"2025-09-27T16:33:02.489999Z","shell.execute_reply.started":"2025-09-27T16:33:02.480177Z","shell.execute_reply":"2025-09-27T16:33:02.489179Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The first two principal components explain about 44% and 20% of the variance respectively, capturing roughly 64% of the total variation in the star features. This indicates that while these components summarize a majority of the data structure, a significant portion of variance remains in higher dimensions. This partial coverage helps explain why stratifying folds based on these components or simple feature clusters is unlikely to produce clearly separated groups, supporting the decision to use random folds instead.","metadata":{}},{"cell_type":"code","source":"# Use KMeans to find clusters in the data, selecting 3 clusters based on domain knowledge/elbow method\nn_clusters = 3\nkmeans = KMeans(n_clusters=n_clusters, random_state=42, n_init=10)\ncluster_labels = kmeans.fit_predict(pca_result)\nstarinfo['cluster'] = cluster_labels","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:53:27.606238Z","iopub.execute_input":"2025-09-27T16:53:27.606643Z","iopub.status.idle":"2025-09-27T16:53:27.671339Z","shell.execute_reply.started":"2025-09-27T16:53:27.606617Z","shell.execute_reply":"2025-09-27T16:53:27.670361Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Visualize clusters in 2D PCA space\nplt.figure(figsize=(8,6))\nfor cluster in range(n_clusters):\n    plt.scatter(\n        pca_result[cluster_labels == cluster, 0],\n        pca_result[cluster_labels == cluster, 1],\n        label=f'Cluster {cluster}', alpha=0.5\n    )\nplt.xlabel('PCA Component 1')\nplt.ylabel('PCA Component 2')\nplt.title('PCA of starinfo features with KMeans Clusters')\nplt.legend()\nplt.grid(True)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:53:28.647469Z","iopub.execute_input":"2025-09-27T16:53:28.64813Z","iopub.status.idle":"2025-09-27T16:53:28.988205Z","shell.execute_reply.started":"2025-09-27T16:53:28.648086Z","shell.execute_reply":"2025-09-27T16:53:28.987098Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The 2D PCA plot with three clusters shows a single large, amorphous blob divided into three loosely defined sections rather than distinct, well-separated clusters. This indicates that natural cluster boundaries are weak or diffuse in the principal component space.\n\nBecause the clusters are not clearly separated, stratifying folds based on these clusters would not create meaningful or stable groups. Instead, this supports the choice to avoid stratification here and rely on random folds to ensure balanced and representative splits.","metadata":{}},{"cell_type":"code","source":"# Silhouette score measures cluster separation quality; higher is better\nfrom sklearn.metrics import silhouette_score\nscore = silhouette_score(pca_result, cluster_labels)\nprint(f\"Silhouette score: {score:.3f}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:35:50.018312Z","iopub.execute_input":"2025-09-27T16:35:50.018681Z","iopub.status.idle":"2025-09-27T16:35:50.075861Z","shell.execute_reply.started":"2025-09-27T16:35:50.018655Z","shell.execute_reply":"2025-09-27T16:35:50.074632Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The silhouette score measures how well samples fit within their assigned clusters compared to other clusters. Scores range from -1 to 1, where higher values indicate better-defined and more separated clusters.\n\nA score of 0.362 indicates weak cluster separation, meaning the clusters overlap substantially and are not very distinct. This aligns with the visual observation of overlapping amorphous cluster regions, further supporting that these clusters do not provide a strong basis for stratification.","metadata":{}},{"cell_type":"markdown","source":"***Experimenting with Alternative Clustering Methods***\n\nWe also try other clustering techniques like Agglomerative Clustering and DBSCAN to explore the data structure from different angles. These can uncover hierarchical or density-based clusters, potentially suggesting meaningful strata.\n\nVisualizing these clusters in PCA space allows us to qualitatively assess cluster separation and overlap.","metadata":{}},{"cell_type":"code","source":"# ## Alternative Clustering Methods Exploration\nfrom sklearn.cluster import AgglomerativeClustering\n\n# Agglomerative clustering with 5 clusters as another method\nagg_model = AgglomerativeClustering(n_clusters=5)\nagg_labels = agg_model.fit_predict(scaled_features)\n\nplt.figure(figsize=(8,6))\nfor lbl in np.unique(agg_labels):\n    plt.scatter(\n        pca_result[agg_labels == lbl, 0],\n        pca_result[agg_labels == lbl, 1],\n        label=f'Cluster {lbl}', alpha=0.6,\n    )\nplt.xlabel('PCA Component 1')\nplt.ylabel('PCA Component 2')\nplt.title('Agglomerative Clustering Visualization')\nplt.legend()\nplt.grid(True)\nplt.show()\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:39:36.472611Z","iopub.execute_input":"2025-09-27T16:39:36.472948Z","iopub.status.idle":"2025-09-27T16:39:36.860348Z","shell.execute_reply.started":"2025-09-27T16:39:36.472921Z","shell.execute_reply":"2025-09-27T16:39:36.859137Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The agglomerative clustering with 5 clusters results in an amorphous blob mostly divided into four visually distinct color groups. The fifth cluster, shown in orange, largely overlaps with the blue cluster, indicating poor separation between these two groups.\n\nThis overlap and lack of clearly separated clusters suggest that the hierarchical clustering does not reveal strong natural groupings in the data. Consequently, using these cluster assignments for stratified folds would not create meaningful or stable splits.","metadata":{}},{"cell_type":"markdown","source":"***3D PCA Visualization of Clusters***\n\nIncreasing the PCA dimensionality to three components lets us explore the data clusters in 3D space, adding more nuance to the cluster shapes and separations.\n\nThis visualization further supports our understanding of the complexity or uniformity in the feature space.","metadata":{}},{"cell_type":"code","source":"# ## 3D PCA Cluster Visualization\n\nfrom mpl_toolkits.mplot3d import Axes3D  # Required for 3D projection\n\n# Perform PCA with 3 components for 3D plotting\npca_3d = PCA(n_components=3)\npca_3d_result = pca_3d.fit_transform(scaled_features)\nprint(f\"PCA Explained Variance Ratio (3D): {pca_3d.explained_variance_ratio_}\")\n\n# Visualize clusters in 3D PCA space\nfig = plt.figure(figsize=(10, 7))\nax = fig.add_subplot(111, projection='3d')\nfor cluster in range(n_clusters):\n    mask = cluster_labels == cluster\n    ax.scatter(\n        pca_3d_result[mask, 0],\n        pca_3d_result[mask, 1],\n        pca_3d_result[mask, 2],\n        label=f\"Cluster {cluster}\", alpha=0.5, s=35\n    )\nax.set_xlabel('PCA Component 1')\nax.set_ylabel('PCA Component 2')\nax.set_zlabel('PCA Component 3')\nax.set_title('3D PCA of starinfo features with KMeans Clusters')\nax.legend()\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:42:00.144193Z","iopub.execute_input":"2025-09-27T16:42:00.145157Z","iopub.status.idle":"2025-09-27T16:42:00.463587Z","shell.execute_reply.started":"2025-09-27T16:42:00.145124Z","shell.execute_reply":"2025-09-27T16:42:00.462386Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The first three principal components explain approximately 44%, 20%, and 14% of the variance respectively, together capturing around 78% of the total data variation. This adds some information beyond the first two components, but still leaves notable variance in higher dimensions.\n\nThe 3D PCA visualization, similar to the 2D case, shows a mostly continuous blob split into three color-coded sections without clear or well-separated clusters. This suggests the absence of strong natural grouping in the main components, reinforcing that stratification based on these clusters would be ineffective.","metadata":{}},{"cell_type":"markdown","source":"***Analyzing Prediction Errors by Star/Planet Parameters***\n\nFinally, we examine how the model's per-planet prediction errors relate to various star and planet features. This analysis helps determine if some features correlate with higher or lower error, which could justify stratified folds.\n\nWe use scatter plots, binned boxplots, and correlation coefficients to detect any notable dependencies or patterns.","metadata":{}},{"cell_type":"code","source":"# ## Analysis of Prediction Errors Relative to Star Features\n\n# Load model predictions and ground truth spectra\npredictions = np.load('/kaggle/input/predictions-0/predictions.npy')\ntrain = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/train.csv')\n\n# Extract spectral columns (wavelength data)\nspectrum_col_names = train.columns.drop('planet_id')\n\n# Construct ground truth spectra matrix ordered by planet_ids\nground_truth = np.vstack([\n    train.loc[train['planet_id'] == pid, spectrum_col_names].values.flatten().astype(float)\n    for pid in planet_ids\n])\n\n# Calculate mean squared error per planet\nmse_per_planet = np.mean((predictions - ground_truth) ** 2, axis=1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:42:43.755013Z","iopub.execute_input":"2025-09-27T16:42:43.755384Z","iopub.status.idle":"2025-09-27T16:42:45.167527Z","shell.execute_reply.started":"2025-09-27T16:42:43.755359Z","shell.execute_reply":"2025-09-27T16:42:45.166361Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load starinfo and select features in correct planet_id order\nfeatures = ['Rs', 'Ms', 'Ts', 'Mp', 'P', 'sma', 'i']\nstarinfo = pd.read_csv('/kaggle/input/ariel-data-challenge-2025/train_star_info.csv')\nif 'planet_id' in starinfo.columns:\n    starinfo = starinfo.set_index('planet_id')\nstar_features = starinfo.loc[planet_ids, features].values","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:43:01.65647Z","iopub.execute_input":"2025-09-27T16:43:01.65678Z","iopub.status.idle":"2025-09-27T16:43:01.668773Z","shell.execute_reply.started":"2025-09-27T16:43:01.656757Z","shell.execute_reply":"2025-09-27T16:43:01.667662Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Plot scatter of each feature against prediction error (MSE)\nfor i, feat in enumerate(features):\n    plt.figure()\n    plt.scatter(starinfo[feat], mse_per_planet, alpha=0.6)\n    plt.xlabel(feat)\n    plt.ylabel(\"Per-planet MSE\")\n    plt.title(f\"Prediction Error vs {feat}\")\n    plt.ylim(top=0.0002)\n    plt.grid(True)\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:43:11.464021Z","iopub.execute_input":"2025-09-27T16:43:11.464389Z","iopub.status.idle":"2025-09-27T16:43:12.758634Z","shell.execute_reply.started":"2025-09-27T16:43:11.464362Z","shell.execute_reply":"2025-09-27T16:43:12.757274Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Binned boxplots of error by quartiles of each feature\ndf = pd.DataFrame({f:starinfo[f] for f in features})\ndf['mse_error'] = mse_per_planet\n\nfor feat in features:\n    df['bin'] = pd.qcut(df[feat], 4, labels=False)\n    plt.figure()\n    sns.boxplot(x=df['bin'], y=df['mse_error'])\n    plt.xlabel(f\"{feat} (binned)\")\n    plt.ylabel(\"Per-planet MSE\")\n    plt.title(f\"Error by {feat} Quartile\")\n    plt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:43:45.049038Z","iopub.execute_input":"2025-09-27T16:43:45.049392Z","iopub.status.idle":"2025-09-27T16:43:46.386364Z","shell.execute_reply.started":"2025-09-27T16:43:45.049367Z","shell.execute_reply":"2025-09-27T16:43:46.38534Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Print correlation coefficients between features and error\nfor i, feat in enumerate(features):\n    corr = np.corrcoef(starinfo[feat], mse_per_planet)[0,1]\n    print(f\"Correlation of {feat} with error: {corr:.3f}\")\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:44:03.29365Z","iopub.execute_input":"2025-09-27T16:44:03.294021Z","iopub.status.idle":"2025-09-27T16:44:03.308547Z","shell.execute_reply.started":"2025-09-27T16:44:03.293993Z","shell.execute_reply":"2025-09-27T16:44:03.307145Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The correlations between prediction error and planetary or stellar features are all very close to zero, indicating no strong relationship. This suggests that prediction errors are fairly uniform across different feature values, supporting the use of random folds instead of stratifying by these parameters.","metadata":{}},{"cell_type":"code","source":"# Additional visualization: Planet mass vs Stellar mass\nplt.figure(figsize=(6, 6))\nplt.scatter(starinfo['Ms'], starinfo['Mp'], alpha=0.6, edgecolor='k')\nplt.xlabel('Stellar Mass (Ms)')\nplt.ylabel('Planet Mass (Mp)')\nplt.title('Planet Mass vs. Stellar Mass')\nplt.grid(True)\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-09-27T16:44:13.214259Z","iopub.execute_input":"2025-09-27T16:44:13.21463Z","iopub.status.idle":"2025-09-27T16:44:13.429842Z","shell.execute_reply.started":"2025-09-27T16:44:13.214604Z","shell.execute_reply":"2025-09-27T16:44:13.428714Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"***Concluding Remarks on Stratification***\n\nBased on the exploratory analyses, clustering results, and error correlation checks, we did not find strong reasons to enforce stratification during fold splitting.\n\nThe random KFold with shuffling provides balanced, well-mixed folds without added complexity. This approach reduces risk of unintended biases or data leakage, and is simpler to implement.\n\nThis notebook demonstrates the due diligence behind this choice, providing transparency and confidence in our validation strategy.","metadata":{}}]}