{
  "id": 609276,
  "title": "6th Place Solution Batman-Minuit",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609276",
  "author_name": "Vicens Gaitan",
  "post_date": "2025-09-25T13:59:48.953000",
  "votes": 28,
  "comment_count": 11,
  "views": 0,
  "content": "<h1>Ariel 25 Challenge: A Three-Stage Pipeline for Robust Transmission Spectra</h1>\n<p>I'd first like to thank the organizers for putting on such an interesting competition. They deserve credit both for their scientific rigor and for the intricate puzzles they used to generate the data—some of which I still haven't managed to solve.</p>\n<h2>1. Summary</h2>\n<p>This solution presents a comprehensive, three-stage pipeline designed to extract high-fidelity transmission spectra from the FGS1 and AIR-CH0 Ariel instrument data. The approach combines physics-based modeling with a machine learning-inspired post-processing framework to tackle the key challenges of instrumental systematics and uncertainty estimation.</p>\n<ol>\n<li><p>Stage 1: GPU-Accelerated Preprocessing. Raw sensor data is converted into clean, calibrated light curves using a CuPy-based pipeline. This stage handles standard instrumental corrections, including non-linearity, dark current, flat-fielding, and jitter, all performed efficiently on the GPU.</p></li>\n<li><p>Stage 2: Hierarchical Transit Fitting. I use the <a href=\"https://lkreidberg.github.io/batman/docs/html/index.html\" target=\"_blank\"><strong>batman</strong></a> transit modeling library and the <a href=\"https://scikit-hep.org/iminuit/\" target=\"_blank\"><strong>iminuit</strong></a> optimizer to perform a multi-step fit. A robust global model is first fit to the combined FGS and AIRS light curves to constrain key orbital parameters. This is followed by a 2D drift correction and a wavelength-by-wavelength fit to extract the initial transmission spectrum.</p></li>\n<li><p>Stage 3: Cross-Validated Post-Processing Ensemble. The initial spectra are refined using an ensemble of models trained with Grouped K-Fold Cross-Validation. This final stage blends a PCA-regularized signal with a smoothed signal and employs a parameterized model to predict the final uncertainties (sigmas), optimizing all parameters directly against the competition score.</p></li>\n</ol>\n<p>This multi-stage design ensures that physical constraints are respected in the initial fit, while the final model has the flexibility to learn and correct for residual systematic errors and produce well-calibrated uncertainties.</p>\n<hr>\n<h2>2. Methodology</h2>\n<h3>Stage 1: Preprocessing Raw Data</h3>\n<p>The first step in the pipeline is to transform the raw sensor readouts into scientifically useful light curves. This entire process is executed on the GPU using CuPy for maximum efficiency.</p>\n<p>Key Preprocessing Steps:</p>\n<p>• Initial Calibrations: The pipeline begins by applying standard instrumental corrections. This includes an ADC correction for gain and offset, a 5th-degree polynomial correction for detector non-linearity (apply_linear_corr_fast), and subtraction of a scaled master dark frame (clean_dark).</p>\n<p>• Flat-Fielding &amp; Masking: A master flat frame is applied to correct for pixel-to-pixel sensitivity variations. Pixels identified as \"hot\" (via sigma-clipping the dark frame) or \"dead\" are masked.</p>\n<p>• Jitter Correction: For the FGS1 sensor, we apply a center-of-mass  regression to correct for flux variations caused by image motion on the detector. The flux is de-correlated from the normalized X and Y  positions of the stellar image. While a PCA-based jitter correction method was developed for the AIRS sensor, it was found to be less effective and is disabled in the final pipeline.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2Fc209358dd02c309ac5067f60b7903724%2Fjitter.png?generation=1758796485996693&amp;alt=media\" alt=\"\"></p>\n<p>• Signal Extraction &amp; Cleaning:</p>\n<ul>\n<li>For FGS1, the final flux is extracted using simple aperture photometry. PSF photometry gives lower SNR signals, probably due to the jitter.</li>\n<li>For AIRS, the signal is extracted from the central detector region, and a background signal, calculated from the top and bottom edges of the detector, is subtracted. This is not optimal because the spectroscopy signal extends i a difractive way to the whole sensor, making dificult to  identify the aditive frequency  dependent background and  making necessary a global  scale factor and an spectrum slope correction in the post processing stage<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F7f6d6424c53a3475488d53772ec8b701%2Fair_psf.png?generation=1758806615604977&amp;alt=media\" alt=\"\"></li>\n<li>A spike-cleaning algorithm  is applied to the final time series to remove cosmic ray hits by identifying and replacing outliers in the frame-to-frame difference.</li>\n</ul>\n<p>• Binning: To improve the signal-to-noise ratio, the final time series for both instruments are binned by averaging consecutive frames.</p>\n<h3>Stage 2: Physics-Based Spectral Fitting</h3>\n<p>With clean light curves for both instruments, we extract the initial transmission spectrum. This process is parallelized using pqdm to efficiently process all observations.</p>\n<p>Hierarchical Fitting Strategy:</p>\n<p><strong>1. Global Combined Fit</strong>: We first fit the FGS light curve and the wavelength-averaged AIRS light curve <strong>simultaneously</strong> using fit_combined_curves. This crucial step provides robust constraints on the shared physical parameters of the system (e.g., orbital period per, inclination inc, transit time t0), which are initialized from the provided star_info.csv file.  The transit is modeled using batman, asuming a <strong>quadratic limb darkening model</strong>, using 2 parameters, and instrumental trends are modeled with a polynomial baseline in time.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F04530cea65be4483a14e6effec93bbe2%2Ffits.png?generation=1758796534186505&amp;alt=media\" alt=\"\"></p>\n<p><strong>2. Refined Fit &amp; Drift Correction</strong>: The results from the first fit are used to initialize a second, refined fit on normalized data. Subsequently, a 2D instrumental drift model is fit to the out-of-transit portion of the AIRS data cube. This model (Drift class) consists of two 4th-order polynomials, one for the frequency and one for the temporal axis, and it corrects for slow-varying systematic patterns across the detector. The data cube is then divided by this drift model.<br>\n<strong>3. Wavelength-by-Wavelength Fit</strong>: The final step is to measure the transit depth in each individual wavelength channel of the drift-corrected AIRS data. The fit_w function iterates through the channels, fitting the data within a small moving window (deltaw=2) to boost signal-to-noise. In this fit, most orbital parameters are fixed to the values from the global fit, and only the transit depth (dipa) and baseline normalization (A0a) are allowed to vary.<br>\nThis process results in the \"raw\" spectrum  and its associated uncertainties , which serve as the primary inputs for our final modeling stage.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F990115e5eb8fa6e302f721f552a1cb94%2Fspectrum.png?generation=1758796611221256&amp;alt=media\" alt=\"\"></p>\n<h2>Stage 3: Post-Processing and Uncertainty Modeling</h2>\n<p>The final stage of the solution is a post-processing model that refines the raw spectra and, critically, learns a robust model for the final uncertainties. This is where the bulk of the \"machine learning\" occurs.</p>\n<p>Model Architecture (build_preds): The model generates a final prediction by blending two different representations of the raw spectrum (preds1):</p>\n<p><strong>PCA-Regularized Prediction (preds0)</strong>: The raw spectrum undergoes a slope correction and is then projected onto a pre-defined PCA basis (components.npy). This de-noises the spectrum and imposes a strong regularization prior, capturing the most common modes of variation.<br>\n<strong>Smoothed Prediction (preds2)</strong>: The raw AIRS spectrum is smoothed using a Savitzky-Golay filter to reduce high-frequency noise while preserving broader spectral features.</p>\n<p>The final mean prediction is a weighted average of these two models: preds = app0 + (1-ap)preds2, where ap is a learned parameter.</p>\n<p><strong>Sigma Model</strong>: Predicting accurate uncertainties is key to maximizing the competition score. I developed a complex, empirically-derived formula to model the final sigma for each data point. This formula combines multiple sources of uncertainty:</p>\n<ul>\n<li>a1* err^e1 : The propagated error from the initial fit, with a learned exponent.</li>\n<li>a01/p1^e2  : A term related to the signal magnitude.</li>\n<li>b0 * sigma0: A term proportional to the raw signal's volatility.</li>\n<li>ae * np.abs(p0-preds2): A term that increases uncertainty where the PCA and smoothed models disagree, capturing model <br>\nuncertainty or unknow spectra.</li>\n</ul>\n<p>Finally ther is a sigma  adjustment  samples with very few or very many out-of-transit points, or very differnt dip por FGS1 and AIR-CH0</p>\n<h1>Training and Ensembling:</h1>\n<p><strong>Objective</strong>: The free parameters of the model (ap, e1, e2, a1, a0, b0, ae, etc.) are optimized using iminuit to directly maximize the competition score.</p>\n<p><strong>Cross-Validation</strong>: To build a robust model that generalizes well, we employ a 10-fold Grouped K-Fold Cross-Validation strategy, using planet_id to group the data. This ensures that all observations of a single planet are kept within the same fold, preventing data leakage.</p>\n<p><strong>Calibration</strong>: Within each fold, after the primary parameters are optimized, we calculate a wavelength-dependent calibration factor (alpha) that scales the predicted sigmas to best match the variance observed in the training data. This factor is then applied to the validation set predictions.</p>\n<p><strong>Ensembling</strong>: The final submission is generated by averaging the calibrated predictions and sigmas from the 10 models trained during the cross-validation process. This ensembling technique reduces variance and improves the final score.</p>\n<p>The entire training process, including the parallel optimization of each fold, is managed by the CV_model function in the model.py file.</p>\n<hr>\n<ol>\n<li>What Did Not Work<br>\nI attempted to post-process the spectrum and sigma values using Gaussian Processes and Ridge Regression, but the results were worse than using the blending of a smoothed/PCA-regularized model and the empirical sigma formula.</li>\n</ol>\n<p>In the final weeks, I spent time trying to isolate the additive foreground noise in the AIRS data for each sample, which would have eliminated the need for the global factor and slope compensation. However, I was unable to find a way to de-correlate signal and noise. I observed a clear structure in the spectroscopic PSF, but all attempts to extract the noise yielded incorrect values that failed to compensate for the systematics.</p>\n<ol>\n<li>A Comment on Efficiency<br>\nThe choice of batman and iMinuit was driven by speed. iMinuit is a Python wrapper for the C++ port of the MINUIT Fortran code (developed by Fred James in the '70s; <a href=\"https://www.sciencedirect.com/science/article/abs/pii/0010465575900399?via%3Dihub\" target=\"_blank\">Paper</a>, <a href=\"https://en.wikipedia.org/wiki/MINUIT\" target=\"_blank\">Wikipedia</a>), and it is typically faster than generic SciPy optimization routines. Additionally, the use of GPU acceleration for data processing via CuPy results in a very fast pipeline because all processes are fully parallelized. The submission process on Kaggle using a P100 instance takes less than 1.5 hours, and on a machine with 100 cores and 4  nVidia L40s GPUs, the full preprocess, fitting, and post-process takes around 5 minutes. The post-process parameter adjustment can also be done in minutes.</li>\n</ol>\n<hr>\n<h2>6. Conclusion</h2>\n<p>This solution demonstrates the power of a hybrid approach, blending physics-informed transit modeling with a flexible, data-driven post-processing framework. By first extracting a reasonable physical spectrum and then refining it with a model trained to optimize the specific competition metric, we can effectively correct for complex instrumental systematics and produce highly accurate, well-calibrated transmission spectra. The use of GPU acceleration, parallel processing, and a robust cross-validation scheme ensures that this complex pipeline remains computationally efficient and resistant to overfitting.</p>",
  "messages": [
    {
      "id": 3294167,
      "postDate": "2025-09-25T13:59:48.953Z",
      "content": "<h1>Ariel 25 Challenge: A Three-Stage Pipeline for Robust Transmission Spectra</h1>\n<p>I'd first like to thank the organizers for putting on such an interesting competition. They deserve credit both for their scientific rigor and for the intricate puzzles they used to generate the data—some of which I still haven't managed to solve.</p>\n<h2>1. Summary</h2>\n<p>This solution presents a comprehensive, three-stage pipeline designed to extract high-fidelity transmission spectra from the FGS1 and AIR-CH0 Ariel instrument data. The approach combines physics-based modeling with a machine learning-inspired post-processing framework to tackle the key challenges of instrumental systematics and uncertainty estimation.</p>\n<ol>\n<li><p>Stage 1: GPU-Accelerated Preprocessing. Raw sensor data is converted into clean, calibrated light curves using a CuPy-based pipeline. This stage handles standard instrumental corrections, including non-linearity, dark current, flat-fielding, and jitter, all performed efficiently on the GPU.</p></li>\n<li><p>Stage 2: Hierarchical Transit Fitting. I use the <a href=\"https://lkreidberg.github.io/batman/docs/html/index.html\" target=\"_blank\"><strong>batman</strong></a> transit modeling library and the <a href=\"https://scikit-hep.org/iminuit/\" target=\"_blank\"><strong>iminuit</strong></a> optimizer to perform a multi-step fit. A robust global model is first fit to the combined FGS and AIRS light curves to constrain key orbital parameters. This is followed by a 2D drift correction and a wavelength-by-wavelength fit to extract the initial transmission spectrum.</p></li>\n<li><p>Stage 3: Cross-Validated Post-Processing Ensemble. The initial spectra are refined using an ensemble of models trained with Grouped K-Fold Cross-Validation. This final stage blends a PCA-regularized signal with a smoothed signal and employs a parameterized model to predict the final uncertainties (sigmas), optimizing all parameters directly against the competition score.</p></li>\n</ol>\n<p>This multi-stage design ensures that physical constraints are respected in the initial fit, while the final model has the flexibility to learn and correct for residual systematic errors and produce well-calibrated uncertainties.</p>\n<hr>\n<h2>2. Methodology</h2>\n<h3>Stage 1: Preprocessing Raw Data</h3>\n<p>The first step in the pipeline is to transform the raw sensor readouts into scientifically useful light curves. This entire process is executed on the GPU using CuPy for maximum efficiency.</p>\n<p>Key Preprocessing Steps:</p>\n<p>• Initial Calibrations: The pipeline begins by applying standard instrumental corrections. This includes an ADC correction for gain and offset, a 5th-degree polynomial correction for detector non-linearity (apply_linear_corr_fast), and subtraction of a scaled master dark frame (clean_dark).</p>\n<p>• Flat-Fielding &amp; Masking: A master flat frame is applied to correct for pixel-to-pixel sensitivity variations. Pixels identified as \"hot\" (via sigma-clipping the dark frame) or \"dead\" are masked.</p>\n<p>• Jitter Correction: For the FGS1 sensor, we apply a center-of-mass  regression to correct for flux variations caused by image motion on the detector. The flux is de-correlated from the normalized X and Y  positions of the stellar image. While a PCA-based jitter correction method was developed for the AIRS sensor, it was found to be less effective and is disabled in the final pipeline.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2Fc209358dd02c309ac5067f60b7903724%2Fjitter.png?generation=1758796485996693&amp;alt=media\" alt=\"\"></p>\n<p>• Signal Extraction &amp; Cleaning:</p>\n<ul>\n<li>For FGS1, the final flux is extracted using simple aperture photometry. PSF photometry gives lower SNR signals, probably due to the jitter.</li>\n<li>For AIRS, the signal is extracted from the central detector region, and a background signal, calculated from the top and bottom edges of the detector, is subtracted. This is not optimal because the spectroscopy signal extends i a difractive way to the whole sensor, making dificult to  identify the aditive frequency  dependent background and  making necessary a global  scale factor and an spectrum slope correction in the post processing stage<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F7f6d6424c53a3475488d53772ec8b701%2Fair_psf.png?generation=1758806615604977&amp;alt=media\" alt=\"\"></li>\n<li>A spike-cleaning algorithm  is applied to the final time series to remove cosmic ray hits by identifying and replacing outliers in the frame-to-frame difference.</li>\n</ul>\n<p>• Binning: To improve the signal-to-noise ratio, the final time series for both instruments are binned by averaging consecutive frames.</p>\n<h3>Stage 2: Physics-Based Spectral Fitting</h3>\n<p>With clean light curves for both instruments, we extract the initial transmission spectrum. This process is parallelized using pqdm to efficiently process all observations.</p>\n<p>Hierarchical Fitting Strategy:</p>\n<p><strong>1. Global Combined Fit</strong>: We first fit the FGS light curve and the wavelength-averaged AIRS light curve <strong>simultaneously</strong> using fit_combined_curves. This crucial step provides robust constraints on the shared physical parameters of the system (e.g., orbital period per, inclination inc, transit time t0), which are initialized from the provided star_info.csv file.  The transit is modeled using batman, asuming a <strong>quadratic limb darkening model</strong>, using 2 parameters, and instrumental trends are modeled with a polynomial baseline in time.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F04530cea65be4483a14e6effec93bbe2%2Ffits.png?generation=1758796534186505&amp;alt=media\" alt=\"\"></p>\n<p><strong>2. Refined Fit &amp; Drift Correction</strong>: The results from the first fit are used to initialize a second, refined fit on normalized data. Subsequently, a 2D instrumental drift model is fit to the out-of-transit portion of the AIRS data cube. This model (Drift class) consists of two 4th-order polynomials, one for the frequency and one for the temporal axis, and it corrects for slow-varying systematic patterns across the detector. The data cube is then divided by this drift model.<br>\n<strong>3. Wavelength-by-Wavelength Fit</strong>: The final step is to measure the transit depth in each individual wavelength channel of the drift-corrected AIRS data. The fit_w function iterates through the channels, fitting the data within a small moving window (deltaw=2) to boost signal-to-noise. In this fit, most orbital parameters are fixed to the values from the global fit, and only the transit depth (dipa) and baseline normalization (A0a) are allowed to vary.<br>\nThis process results in the \"raw\" spectrum  and its associated uncertainties , which serve as the primary inputs for our final modeling stage.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F990115e5eb8fa6e302f721f552a1cb94%2Fspectrum.png?generation=1758796611221256&amp;alt=media\" alt=\"\"></p>\n<h2>Stage 3: Post-Processing and Uncertainty Modeling</h2>\n<p>The final stage of the solution is a post-processing model that refines the raw spectra and, critically, learns a robust model for the final uncertainties. This is where the bulk of the \"machine learning\" occurs.</p>\n<p>Model Architecture (build_preds): The model generates a final prediction by blending two different representations of the raw spectrum (preds1):</p>\n<p><strong>PCA-Regularized Prediction (preds0)</strong>: The raw spectrum undergoes a slope correction and is then projected onto a pre-defined PCA basis (components.npy). This de-noises the spectrum and imposes a strong regularization prior, capturing the most common modes of variation.<br>\n<strong>Smoothed Prediction (preds2)</strong>: The raw AIRS spectrum is smoothed using a Savitzky-Golay filter to reduce high-frequency noise while preserving broader spectral features.</p>\n<p>The final mean prediction is a weighted average of these two models: preds = app0 + (1-ap)preds2, where ap is a learned parameter.</p>\n<p><strong>Sigma Model</strong>: Predicting accurate uncertainties is key to maximizing the competition score. I developed a complex, empirically-derived formula to model the final sigma for each data point. This formula combines multiple sources of uncertainty:</p>\n<ul>\n<li>a1* err^e1 : The propagated error from the initial fit, with a learned exponent.</li>\n<li>a01/p1^e2  : A term related to the signal magnitude.</li>\n<li>b0 * sigma0: A term proportional to the raw signal's volatility.</li>\n<li>ae * np.abs(p0-preds2): A term that increases uncertainty where the PCA and smoothed models disagree, capturing model <br>\nuncertainty or unknow spectra.</li>\n</ul>\n<p>Finally ther is a sigma  adjustment  samples with very few or very many out-of-transit points, or very differnt dip por FGS1 and AIR-CH0</p>\n<h1>Training and Ensembling:</h1>\n<p><strong>Objective</strong>: The free parameters of the model (ap, e1, e2, a1, a0, b0, ae, etc.) are optimized using iminuit to directly maximize the competition score.</p>\n<p><strong>Cross-Validation</strong>: To build a robust model that generalizes well, we employ a 10-fold Grouped K-Fold Cross-Validation strategy, using planet_id to group the data. This ensures that all observations of a single planet are kept within the same fold, preventing data leakage.</p>\n<p><strong>Calibration</strong>: Within each fold, after the primary parameters are optimized, we calculate a wavelength-dependent calibration factor (alpha) that scales the predicted sigmas to best match the variance observed in the training data. This factor is then applied to the validation set predictions.</p>\n<p><strong>Ensembling</strong>: The final submission is generated by averaging the calibrated predictions and sigmas from the 10 models trained during the cross-validation process. This ensembling technique reduces variance and improves the final score.</p>\n<p>The entire training process, including the parallel optimization of each fold, is managed by the CV_model function in the model.py file.</p>\n<hr>\n<ol>\n<li>What Did Not Work<br>\nI attempted to post-process the spectrum and sigma values using Gaussian Processes and Ridge Regression, but the results were worse than using the blending of a smoothed/PCA-regularized model and the empirical sigma formula.</li>\n</ol>\n<p>In the final weeks, I spent time trying to isolate the additive foreground noise in the AIRS data for each sample, which would have eliminated the need for the global factor and slope compensation. However, I was unable to find a way to de-correlate signal and noise. I observed a clear structure in the spectroscopic PSF, but all attempts to extract the noise yielded incorrect values that failed to compensate for the systematics.</p>\n<ol>\n<li>A Comment on Efficiency<br>\nThe choice of batman and iMinuit was driven by speed. iMinuit is a Python wrapper for the C++ port of the MINUIT Fortran code (developed by Fred James in the '70s; <a href=\"https://www.sciencedirect.com/science/article/abs/pii/0010465575900399?via%3Dihub\" target=\"_blank\">Paper</a>, <a href=\"https://en.wikipedia.org/wiki/MINUIT\" target=\"_blank\">Wikipedia</a>), and it is typically faster than generic SciPy optimization routines. Additionally, the use of GPU acceleration for data processing via CuPy results in a very fast pipeline because all processes are fully parallelized. The submission process on Kaggle using a P100 instance takes less than 1.5 hours, and on a machine with 100 cores and 4  nVidia L40s GPUs, the full preprocess, fitting, and post-process takes around 5 minutes. The post-process parameter adjustment can also be done in minutes.</li>\n</ol>\n<hr>\n<h2>6. Conclusion</h2>\n<p>This solution demonstrates the power of a hybrid approach, blending physics-informed transit modeling with a flexible, data-driven post-processing framework. By first extracting a reasonable physical spectrum and then refining it with a model trained to optimize the specific competition metric, we can effectively correct for complex instrumental systematics and produce highly accurate, well-calibrated transmission spectra. The use of GPU acceleration, parallel processing, and a robust cross-validation scheme ensures that this complex pipeline remains computationally efficient and resistant to overfitting.</p>",
      "rawMarkdown": "# Ariel 25 Challenge: A Three-Stage Pipeline for Robust Transmission Spectra\n\nI'd first like to thank the organizers for putting on such an interesting competition. They deserve credit both for their scientific rigor and for the intricate puzzles they used to generate the data—some of which I still haven't managed to solve.\n\n##1. Summary\n\nThis solution presents a comprehensive, three-stage pipeline designed to extract high-fidelity transmission spectra from the FGS1 and AIR-CH0 Ariel instrument data. The approach combines physics-based modeling with a machine learning-inspired post-processing framework to tackle the key challenges of instrumental systematics and uncertainty estimation.\n\n1. Stage 1: GPU-Accelerated Preprocessing. Raw sensor data is converted into clean, calibrated light curves using a CuPy-based pipeline. This stage handles standard instrumental corrections, including non-linearity, dark current, flat-fielding, and jitter, all performed efficiently on the GPU.\n\n2. Stage 2: Hierarchical Transit Fitting. I use the [**batman**](https://lkreidberg.github.io/batman/docs/html/index.html) transit modeling library and the [**iminuit**](https://scikit-hep.org/iminuit/) optimizer to perform a multi-step fit. A robust global model is first fit to the combined FGS and AIRS light curves to constrain key orbital parameters. This is followed by a 2D drift correction and a wavelength-by-wavelength fit to extract the initial transmission spectrum.\n\n3. Stage 3: Cross-Validated Post-Processing Ensemble. The initial spectra are refined using an ensemble of models trained with Grouped K-Fold Cross-Validation. This final stage blends a PCA-regularized signal with a smoothed signal and employs a parameterized model to predict the final uncertainties (sigmas), optimizing all parameters directly against the competition score.\n\n\nThis multi-stage design ensures that physical constraints are respected in the initial fit, while the final model has the flexibility to learn and correct for residual systematic errors and produce well-calibrated uncertainties.\n\n--------------------------------------------------------------------------------\n## 2. Methodology\n\n###Stage 1: Preprocessing Raw Data\nThe first step in the pipeline is to transform the raw sensor readouts into scientifically useful light curves. This entire process is executed on the GPU using CuPy for maximum efficiency.\n\nKey Preprocessing Steps:\n\n• Initial Calibrations: The pipeline begins by applying standard instrumental corrections. This includes an ADC correction for gain and offset, a 5th-degree polynomial correction for detector non-linearity (apply_linear_corr_fast), and subtraction of a scaled master dark frame (clean_dark).\n\n• Flat-Fielding & Masking: A master flat frame is applied to correct for pixel-to-pixel sensitivity variations. Pixels identified as \"hot\" (via sigma-clipping the dark frame) or \"dead\" are masked.\n\n• Jitter Correction: For the FGS1 sensor, we apply a center-of-mass  regression to correct for flux variations caused by image motion on the detector. The flux is de-correlated from the normalized X and Y  positions of the stellar image. While a PCA-based jitter correction method was developed for the AIRS sensor, it was found to be less effective and is disabled in the final pipeline.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2Fc209358dd02c309ac5067f60b7903724%2Fjitter.png?generation=1758796485996693&alt=media)\n\n\n\n• Signal Extraction & Cleaning:\n\n   \n* For FGS1, the final flux is extracted using simple aperture photometry. PSF photometry gives lower SNR signals, probably due to the jitter.\n* For AIRS, the signal is extracted from the central detector region, and a background signal, calculated from the top and bottom edges of the detector, is subtracted. This is not optimal because the spectroscopy signal extends i a difractive way to the whole sensor, making dificult to  identify the aditive frequency  dependent background and  making necessary a global  scale factor and an spectrum slope correction in the post processing stage\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F7f6d6424c53a3475488d53772ec8b701%2Fair_psf.png?generation=1758806615604977&alt=media)\n* A spike-cleaning algorithm  is applied to the final time series to remove cosmic ray hits by identifying and replacing outliers in the frame-to-frame difference.\n\n    \n• Binning: To improve the signal-to-noise ratio, the final time series for both instruments are binned by averaging consecutive frames.\n\n###Stage 2: Physics-Based Spectral Fitting\n\nWith clean light curves for both instruments, we extract the initial transmission spectrum. This process is parallelized using pqdm to efficiently process all observations.\n\nHierarchical Fitting Strategy:\n\n**1. Global Combined Fit**: We first fit the FGS light curve and the wavelength-averaged AIRS light curve **simultaneously** using fit_combined_curves. This crucial step provides robust constraints on the shared physical parameters of the system (e.g., orbital period per, inclination inc, transit time t0), which are initialized from the provided star_info.csv file.  The transit is modeled using batman, asuming a **quadratic limb darkening model**, using 2 parameters, and instrumental trends are modeled with a polynomial baseline in time.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F04530cea65be4483a14e6effec93bbe2%2Ffits.png?generation=1758796534186505&alt=media)\n\n**2. Refined Fit & Drift Correction**: The results from the first fit are used to initialize a second, refined fit on normalized data. Subsequently, a 2D instrumental drift model is fit to the out-of-transit portion of the AIRS data cube. This model (Drift class) consists of two 4th-order polynomials, one for the frequency and one for the temporal axis, and it corrects for slow-varying systematic patterns across the detector. The data cube is then divided by this drift model.\n**3. Wavelength-by-Wavelength Fit**: The final step is to measure the transit depth in each individual wavelength channel of the drift-corrected AIRS data. The fit_w function iterates through the channels, fitting the data within a small moving window (deltaw=2) to boost signal-to-noise. In this fit, most orbital parameters are fixed to the values from the global fit, and only the transit depth (dipa) and baseline normalization (A0a) are allowed to vary.\nThis process results in the \"raw\" spectrum  and its associated uncertainties , which serve as the primary inputs for our final modeling stage.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F990115e5eb8fa6e302f721f552a1cb94%2Fspectrum.png?generation=1758796611221256&alt=media)\n\n## Stage 3: Post-Processing and Uncertainty Modeling\n\nThe final stage of the solution is a post-processing model that refines the raw spectra and, critically, learns a robust model for the final uncertainties. This is where the bulk of the \"machine learning\" occurs.\n\nModel Architecture (build_preds): The model generates a final prediction by blending two different representations of the raw spectrum (preds1):\n\n**PCA-Regularized Prediction (preds0)**: The raw spectrum undergoes a slope correction and is then projected onto a pre-defined PCA basis (components.npy). This de-noises the spectrum and imposes a strong regularization prior, capturing the most common modes of variation.\n**Smoothed Prediction (preds2)**: The raw AIRS spectrum is smoothed using a Savitzky-Golay filter to reduce high-frequency noise while preserving broader spectral features.\n\nThe final mean prediction is a weighted average of these two models: preds = app0 + (1-ap)preds2, where ap is a learned parameter.\n\n**Sigma Model**: Predicting accurate uncertainties is key to maximizing the competition score. I developed a complex, empirically-derived formula to model the final sigma for each data point. This formula combines multiple sources of uncertainty:\n\n \n* a1* err^e1 : The propagated error from the initial fit, with a learned exponent.\n* a01/p1^e2  : A term related to the signal magnitude.\n* b0 * sigma0: A term proportional to the raw signal's volatility.\n* ae * np.abs(p0-preds2): A term that increases uncertainty where the PCA and smoothed models disagree, capturing model \nuncertainty or unknow spectra.\n\nFinally ther is a sigma  adjustment  samples with very few or very many out-of-transit points, or very differnt dip por FGS1 and AIR-CH0\n\n#Training and Ensembling:\n\n  **Objective**: The free parameters of the model (ap, e1, e2, a1, a0, b0, ae, etc.) are optimized using iminuit to directly maximize the competition score.\n\n  **Cross-Validation**: To build a robust model that generalizes well, we employ a 10-fold Grouped K-Fold Cross-Validation strategy, using planet_id to group the data. This ensures that all observations of a single planet are kept within the same fold, preventing data leakage.\n\n  **Calibration**: Within each fold, after the primary parameters are optimized, we calculate a wavelength-dependent calibration factor (alpha) that scales the predicted sigmas to best match the variance observed in the training data. This factor is then applied to the validation set predictions.\n  \n   **Ensembling**: The final submission is generated by averaging the calibrated predictions and sigmas from the 10 models trained during the cross-validation process. This ensembling technique reduces variance and improves the final score.\n\nThe entire training process, including the parallel optimization of each fold, is managed by the CV_model function in the model.py file.\n\n--------------------------------------------------------------------------------\n4. What Did Not Work\nI attempted to post-process the spectrum and sigma values using Gaussian Processes and Ridge Regression, but the results were worse than using the blending of a smoothed/PCA-regularized model and the empirical sigma formula.\n\nIn the final weeks, I spent time trying to isolate the additive foreground noise in the AIRS data for each sample, which would have eliminated the need for the global factor and slope compensation. However, I was unable to find a way to de-correlate signal and noise. I observed a clear structure in the spectroscopic PSF, but all attempts to extract the noise yielded incorrect values that failed to compensate for the systematics.\n\n5. A Comment on Efficiency\nThe choice of batman and iMinuit was driven by speed. iMinuit is a Python wrapper for the C++ port of the MINUIT Fortran code (developed by Fred James in the '70s; [Paper](https://www.sciencedirect.com/science/article/abs/pii/0010465575900399?via%3Dihub), [Wikipedia](https://en.wikipedia.org/wiki/MINUIT)), and it is typically faster than generic SciPy optimization routines. Additionally, the use of GPU acceleration for data processing via CuPy results in a very fast pipeline because all processes are fully parallelized. The submission process on Kaggle using a P100 instance takes less than 1.5 hours, and on a machine with 100 cores and 4  nVidia L40s GPUs, the full preprocess, fitting, and post-process takes around 5 minutes. The post-process parameter adjustment can also be done in minutes.\n\n--------------------------------------------------------------------------------\n##6. Conclusion\nThis solution demonstrates the power of a hybrid approach, blending physics-informed transit modeling with a flexible, data-driven post-processing framework. By first extracting a reasonable physical spectrum and then refining it with a model trained to optimize the specific competition metric, we can effectively correct for complex instrumental systematics and produce highly accurate, well-calibrated transmission spectra. The use of GPU acceleration, parallel processing, and a robust cross-validation scheme ensures that this complex pipeline remains computationally efficient and resistant to overfitting.",
      "votes": 28
    },
    {
      "id": 3295343,
      "postDate": "2025-09-28T13:31:56.957Z",
      "content": "<p>Great Work.@Vicens Gaitan</p>",
      "rawMarkdown": "Great Work.@Vicens Gaitan",
      "votes": -1
    },
    {
      "id": 3295118,
      "postDate": "2025-09-27T18:09:21.427Z",
      "content": "<p>Great Work. I appreciate 🥳</p>",
      "rawMarkdown": "Great Work. I appreciate 🥳"
    },
    {
      "id": 3294493,
      "postDate": "2025-09-26T09:17:21.213Z",
      "content": "<p>Nice summary!<br>\nI also used a Minuit (migrad + hesse) + Batman combination. The L-BFGS-B underperformed significantly in my case. However, I used the stellar parameters not only for initialization but also to constrain some of the parameters with Gaussian priors. Did you use the stellar information just for initialization, or did you also apply it as constraints?</p>",
      "rawMarkdown": "Nice summary!\nI also used a Minuit (migrad + hesse) + Batman combination. The L-BFGS-B underperformed significantly in my case. However, I used the stellar parameters not only for initialization but also to constrain some of the parameters with Gaussian priors. Did you use the stellar information just for initialization, or did you also apply it as constraints?",
      "replies": [
        {
          "id": 3294504,
          "postDate": "2025-09-26T09:35:33.377Z",
          "content": "<p>Only for initialization,  but it was important to limit the variation range for each of them based on physical considerations.</p>",
          "rawMarkdown": "Only for initialization,  but it was important to limit the variation range for each of them based on physical considerations.",
          "votes": 1,
          "replies": [
            {
              "id": 3294510,
              "postDate": "2025-09-26T09:47:12.930Z",
              "content": "<p>Thanks for the clarification! I asked it because I'm trying to figure out whether the private star info data was the reason my score dropped so drastically :) </p>",
              "rawMarkdown": "Thanks for the clarification! I asked it because I'm trying to figure out whether the private star info data was the reason my score dropped so drastically :) "
            },
            {
              "id": 3295339,
              "postDate": "2025-09-28T13:18:19.703Z",
              "content": "<p>Just curious, did you fit the <code>e</code> and <code>w</code> parameters (eccentricity and longitude of periastron), or are their values fixed in global combined fit?</p>",
              "rawMarkdown": "Just curious, did you fit the `e` and `w` parameters (eccentricity and longitude of periastron), or are their values fixed in global combined fit?"
            },
            {
              "id": 3295404,
              "postDate": "2025-09-28T16:53:25.667Z",
              "content": "<p>I only fit the orbital parametres in the combined fit and fix it  for the 282 airs  frequency fits. You can take a look at the code  <a href=\"https://www.kaggle.com/code/vicensgaitan/6th-place-solution-pipeline\" target=\"_blank\">here</a> in he fit2.py module </p>\n<p>`</p>\n<pre><code>def fit_w(air, minuit_result, fixed_list, dipa_mean, =1, =2):\n\n\n\n\n\n    initial_guesses = DEFAULT_PARAMS_w.copy()\n    default_limits = DEFAULT_LIMITS.copy()\n\n    # Use results  the combined fit as initial guesses\n     p  initial_guesses:\n         p   fixed_list:\n            initial_guesses[p] = minuit_result.values[p]\n\n    # Fix all parameters except depth ()  normalization ()\n    params_to_fix = list(initial_guesses.keys())\n    params_to_fix.()\n    params_to_fix.()\n</code></pre>\n<p>`</p>",
              "rawMarkdown": "I only fit the orbital parametres in the combined fit and fix it  for the 282 airs  frequency fits. You can take a look at the code  [here](https://www.kaggle.com/code/vicensgaitan/6th-place-solution-pipeline) in he fit2.py module \n\n`\n\n\n    def fit_w(air, minuit_result, fixed_list, dipa_mean, mode=1, delta_w=2):\n    #Fits the transit depth for each wavelength channel ('w') of the AIRS data.\n    #To improve S/N, it fits the mean flux in a moving window of 2 * delta_w\n    #channels around each target channel. Most transit parameters are fixed to the\n    #values from the combined fit (`minuit_result`).\n    \n        initial_guesses = DEFAULT_PARAMS_w.copy()\n        default_limits = DEFAULT_LIMITS.copy()\n\n        # Use results from the combined fit as initial guesses\n        for p in initial_guesses:\n            if p not in fixed_list:\n                initial_guesses[p] = minuit_result.values[p]\n    \n        # Fix all parameters except depth ('dipa') and normalization ('A0a')\n        params_to_fix = list(initial_guesses.keys())\n        params_to_fix.remove('dipa')\n        params_to_fix.remove('A0a')\n`"
            }
          ]
        }
      ]
    },
    {
      "id": 3294274,
      "postDate": "2025-09-25T18:10:51.700Z",
      "content": "<p>Thanks for posting this write-up!  I like the approach, and the speed you're seeing for the wavelength-by-wavelength fit is amazing -- it strikes me as a really impressive achievement, so congratulations on that in particular (as well as the overall success of your analysis of course!).  I was also using a wavelength-by-wavelength fit in my solution, but I struggled a lot with speed (perhaps because I also allowed free parameters for the ingress/egress duration and limb darkening for each wavelength).  </p>\n<p>Mind if I ask a question about your experience setting up this analysis?  One of the challenges I ran into (right at the end, unfortunately, so I didn't manage to fix it) was a drop in the performance of the minimizers I was using when I tried to port my fitting code to the kaggle notebook environment.  It looked like the same call to scipy.optimize.minimize() just…didn't find as good a minimum in the kaggle notebook as it did in my offline code.  Same numpy/scipy/torch versions online and offline, same input data as far as I could see, same arguments to the function invocation, same gradients (at least, as far as I could spot in the first call to the chisquare function), just….different results, as though the fitter just decided to quit early.    I am wondering if you saw anything similar, and if so, how did you resolve it?  I would love to know what to check for next time.  😅</p>",
      "rawMarkdown": "Thanks for posting this write-up!  I like the approach, and the speed you're seeing for the wavelength-by-wavelength fit is amazing -- it strikes me as a really impressive achievement, so congratulations on that in particular (as well as the overall success of your analysis of course!).  I was also using a wavelength-by-wavelength fit in my solution, but I struggled a lot with speed (perhaps because I also allowed free parameters for the ingress/egress duration and limb darkening for each wavelength).  \n\nMind if I ask a question about your experience setting up this analysis?  One of the challenges I ran into (right at the end, unfortunately, so I didn't manage to fix it) was a drop in the performance of the minimizers I was using when I tried to port my fitting code to the kaggle notebook environment.  It looked like the same call to scipy.optimize.minimize() just...didn't find as good a minimum in the kaggle notebook as it did in my offline code.  Same numpy/scipy/torch versions online and offline, same input data as far as I could see, same arguments to the function invocation, same gradients (at least, as far as I could spot in the first call to the chisquare function), just....different results, as though the fitter just decided to quit early.    I am wondering if you saw anything similar, and if so, how did you resolve it?  I would love to know what to check for next time.  😅",
      "replies": [
        {
          "id": 3294437,
          "postDate": "2025-09-26T07:21:16.383Z",
          "content": "<p>Thanks so much for your feedback! 😊</p>\n<p>In my experience, the results from Scipy optimizers (and also in the case of Minuit) are reproducible across different platforms if both the optimizer's parameters and the data types of the function to be optimized and its parameters are kept the same. Although the optimization process is essentially sequential, some of the intermediate calculations might use a parallel backend (for example, OpenBLAS to compute the Hessian). In this specific scenario, we can't guarantee the exact same result, though it shouldn't differ too significantly. Perhaps this isn't the root of the issue, however.</p>\n<p>In my specific case, using Batman to model the transit, it's been critical to impose physical limits on the orbital and limb darkening values. This prevents the optimizer from getting lost and attempting to use the degrees of freedom from the polynomial drift to explain phenomena that actually depend on the orbital parameters.</p>",
          "rawMarkdown": "Thanks so much for your feedback! 😊\n\nIn my experience, the results from Scipy optimizers (and also in the case of Minuit) are reproducible across different platforms if both the optimizer's parameters and the data types of the function to be optimized and its parameters are kept the same. Although the optimization process is essentially sequential, some of the intermediate calculations might use a parallel backend (for example, OpenBLAS to compute the Hessian). In this specific scenario, we can't guarantee the exact same result, though it shouldn't differ too significantly. Perhaps this isn't the root of the issue, however.\n\nIn my specific case, using Batman to model the transit, it's been critical to impose physical limits on the orbital and limb darkening values. This prevents the optimizer from getting lost and attempting to use the degrees of freedom from the polynomial drift to explain phenomena that actually depend on the orbital parameters.",
          "replies": [
            {
              "id": 3295521,
              "postDate": "2025-09-29T02:38:01.913Z",
              "content": "<p>Thanks!  There are some good leads here -- I do notice that numpy.show_config() reports a different blas backend on my desktop compared to the kaggle environment; I haven't yet tried switching my offline code to the one kaggle uses, but it's possible that's what's going on here.  I do have limits on the parameters that govern my signal shape, but not on the polynomial background.  I can play around with that a bit, too.</p>\n<p>But I'll definitely have to give iMinuit a shot; I've used the version in ROOT a lot in the past, but I switched to scipy for this contest because I didn't want to try installing ROOT in a Kaggle notebook.  Glad to know that Minuit is available as a pip install now!</p>",
              "rawMarkdown": "Thanks!  There are some good leads here -- I do notice that numpy.show_config() reports a different blas backend on my desktop compared to the kaggle environment; I haven't yet tried switching my offline code to the one kaggle uses, but it's possible that's what's going on here.  I do have limits on the parameters that govern my signal shape, but not on the polynomial background.  I can play around with that a bit, too.\n\nBut I'll definitely have to give iMinuit a shot; I've used the version in ROOT a lot in the past, but I switched to scipy for this contest because I didn't want to try installing ROOT in a Kaggle notebook.  Glad to know that Minuit is available as a pip install now!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3294174,
      "postDate": "2025-09-25T14:26:37.560Z",
      "content": "<h1>UFC Fight Prediction Analysis</h1>\n<h2>Table of Contents</h2>\n<ol>\n<li>Required Libraries</li>\n<li>Data Loading and Preprocessing</li>\n<li>Exploratory Data Analysis</li>\n<li>Feature Engineering</li>\n<li>Model Implementation (Decision Tree, XGBoost, LightGBM)</li>\n<li>Model Comparison and Evaluation</li>\n<li>Advanced Visualization and Analysis (SHAP, Lift/Calibration, Partial Dependence)</li>\n<li>Results and Conclusions</li>\n<li>Next steps and Caveats</li>\n</ol>\n<hr>\n<blockquote>\n  <p><strong>Notes:</strong></p>\n  <ul>\n  <li>This document is a drop-in Python script / notebook. Replace <code>data_path</code> with your dataset path. The code assumes a tabular fight-level dataset where each row is a fight and available columns include fighter-level stats (e.g. <code>red_*</code>, <code>blue_*</code> or <code>fighter1_*</code>/<code>fighter2_*</code>), and an outcome column like <code>winner</code> or <code>result</code>. Adjust column names as needed.</li>\n  </ul>\n</blockquote>\n<hr>\n<h1>1) Required Libraries</h1>\n<pre><code>\n os\n numpy  np\n pandas  pd\n sklearn.model_selection  train_test_split, StratifiedKFold, cross_val_score, GridSearchCV\n sklearn.preprocessing  StandardScaler, OneHotEncoder\n sklearn.compose  ColumnTransformer\n sklearn.pipeline  Pipeline\n\n\n sklearn.tree  DecisionTreeClassifier\n sklearn.metrics  accuracy_score, precision_score, recall_score, f1_score, roc_auc_score, confusion_matrix, classification_report, log_loss\n xgboost  xgb\n lightgbm  lgb\n\n\n shap\n matplotlib.pyplot  plt\n seaborn  sns\n\n\nRANDOM_STATE = \n\n\n warnings\nwarnings.filterwarnings()\n</code></pre>\n<h1>2) Data Loading and Preprocessing</h1>\n<pre><code>\ndata_path =   \ndf = pd.read_csv(data_path)\n\n(df.shape)\n(df.columns.tolist()[:])\n\n\n   df.columns:\n    df[] = (df[]..lower()..contains()).astype()\n   df.columns:\n    \n    df[] = df[].({:, :})\n:\n     ValueError()\n\n\n()\n(df.isna().().sort_values(ascending=).head())\n\n\n\nto_drop = [, , , ]\n c  to_drop:\n     c  df.columns:\n        df.drop(columns=c, inplace=)\n\n\nnum_cols = df.select_dtypes(include=[,]).columns.tolist()\n\nnum_cols = [c  c  num_cols  c != ]\ncat_cols = df.select_dtypes(include=[, ]).columns.tolist()\n\n(, num_cols[:])\n(, cat_cols[:])\n\n\n c  num_cols:\n    df[c] = df[c].fillna(df[c].median())\n c  cat_cols:\n    df[c] = df[c].fillna()\n\n\ntrain_df, test_df = train_test_split(df, test_size=, stratify=df[], random_state=RANDOM_STATE)\n\nX_train = train_df.drop(columns=[])\ny_train = train_df[]\nX_test = test_df.drop(columns=[])\ny_test = test_df[]\n\n(, X_train.shape, X_test.shape)\n</code></pre>\n<h1>3) Exploratory Data Analysis (brief)</h1>\n<pre><code>\n(y_train.value_counts(normalize=))\n\n\nplt.figure(figsize=(,))\nnum_subset = num_cols[:]  \nsns.heatmap(train_df[num_subset + []].corr(), annot=, cmap=)\nplt.title()\nplt.show()\n\n\n c  num_cols[:]:\n    plt.figure(figsize=(,))\n    sns.boxplot(x=train_df[], y=train_df[c])\n    plt.title()\n    plt.show()\n\n\n c  cat_cols:\n    (c, train_df[c].nunique())\n     train_df[c].nunique() &lt; :\n        display(train_df.groupby(c)[].mean().sort_values())\n</code></pre>\n<h1>4) Feature Engineering</h1>\n<pre><code>\n\npairs = []\n col  df.columns:\n     col.startswith():\n        counterpart =  + col[():]\n         counterpart  df.columns:\n            pairs.append((col, counterpart))\n\n r,b  pairs:\n    newcol = r.replace(,)\n    df[newcol] = df[r] - df[b]\n\n\ntrain_df, test_df = train_test_split(df, test_size=, stratify=df[], random_state=RANDOM_STATE)\nX_train = train_df.drop(columns=[])\ny_train = train_df[]\nX_test = test_df.drop(columns=[])\ny_test = test_df[]\n\n\nall_num = X_train.select_dtypes(include=[,]).columns.tolist()\nall_cat = X_train.select_dtypes(include=[,]).columns.tolist()\n(, (all_num))\n(, (all_cat))\n</code></pre>\n<h1>5) Model Implementation</h1>\n<p>We'll build three pipelines: Decision Tree (sklearn), XGBoost, LightGBM. We'll use a ColumnTransformer to scale numeric features and one-hot encode low-cardinality categoricals.</p>\n<pre><code>\nnumeric_transformer = Pipeline(steps=[(, StandardScaler())])\n\n\nlow_cardinality = [c  c  all_cat  X_train[c].nunique() &lt; ]\nhigh_cardinality = [c  c  all_cat  X_train[c].nunique() &gt;= ]\n\ncategorical_transformer = Pipeline(steps=[(, OneHotEncoder(handle_unknown=, sparse=))])\n\npreprocessor = ColumnTransformer(transformers=[\n    (, numeric_transformer, all_num),\n    (, categorical_transformer, low_cardinality)\n], remainder=)\n\n\n sklearn.model_selection  cross_validate\n\n ():\n    scoring = [,,,,]\n    cv_res = cross_validate(pipe, X, y, cv=cv, scoring=scoring, return_train_score=)\n     {k: np.mean(v)  k,v  cv_res.items()}\n\n\ndt = DecisionTreeClassifier(random_state=RANDOM_STATE)\npipe_dt = Pipeline(steps=[(, preprocessor), (, dt)])\n\n\n(, evaluate_model(pipe_dt, X_train, y_train, cv=))\n\n\nxgb_clf = xgb.XGBClassifier(use_label_encoder=, eval_metric=, random_state=RANDOM_STATE)\npipe_xgb = Pipeline(steps=[(, preprocessor), (, xgb_clf)])\n(, evaluate_model(pipe_xgb, X_train, y_train, cv=))\n\n\nlgb_clf = lgb.LGBMClassifier(random_state=RANDOM_STATE)\npipe_lgb = Pipeline(steps=[(, preprocessor), (, lgb_clf)])\n(, evaluate_model(pipe_lgb, X_train, y_train, cv=))\n</code></pre>\n<h2>Hyperparameter tuning (example grids)</h2>\n<pre><code>\ndt_param_grid = {\n    : [,,,],\n    : [,,]\n}\n\nxgb_param_grid = {\n    : [,],\n    : [,],\n    : [, ]\n}\n\nlgb_param_grid = {\n    : [,],\n    : [, ],\n    : [, ]\n}\n\n\ngs_lgb = GridSearchCV(pipe_lgb, lgb_param_grid, cv=, scoring=, n_jobs=-, verbose=)\ngs_lgb.fit(X_train, y_train)\n(, gs_lgb.best_params_)\n(, gs_lgb.best_score_)\n\n\nbest_lgb = gs_lgb.best_estimator_\n\n\n</code></pre>\n<h1>6) Model Comparison and Evaluation (holdout test)</h1>\n<pre><code>models = {\n    : pipe_dt.fit(X_train, y_train),\n    : pipe_xgb.fit(X_train, y_train),\n    : best_lgb    ()  pipe_lgb.fit(X_train, y_train)\n}\n\nresults = {}\n name, model  models.items():\n    y_pred = model.predict(X_test)\n    y_proba = model.predict_proba(X_test)[:,]\n    results[name] = {\n        : accuracy_score(y_test, y_pred),\n        : precision_score(y_test, y_pred),\n        : recall_score(y_test, y_pred),\n        : f1_score(y_test, y_pred),\n        : roc_auc_score(y_test, y_proba),\n        : log_loss(y_test, y_proba)\n    }\n\npd.DataFrame(results).T.sort_values(, ascending=)\n\n\nbest_name = (results.keys(), key= k: results[k][])\n(, best_name)\nbest_model = models[best_name]\n\ny_pred = best_model.predict(X_test)\ncm = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(,))\nsns.heatmap(cm, annot=, fmt=)\nplt.title()\nplt.xlabel()\nplt.ylabel()\nplt.show()\n</code></pre>\n<h1>7) Advanced Visualization and Analysis</h1>\n<pre><code>\n\n\npre = best_model.named_steps[]\nclf = best_model.named_steps[]\n\n\npre.fit(X_train)\n\nnum_names = all_num\ncat_encoder = pre.named_transformers_.get()\n cat_encoder   :\n    cat_names = (cat_encoder.named_steps[].get_feature_names_out(low_cardinality))\n:\n    cat_names = []\nfeature_names = num_names + cat_names\n\n\nX_test_trans = pre.transform(X_test)\n\n\nexplainer = shap.TreeExplainer(clf)\nshap_values = explainer.shap_values(X_test_trans)\n\n\n (shap_values, ):\n    sv = shap_values[]\n:\n    sv = shap_values\n\n\nshap.summary_plot(sv, features=X_test_trans, feature_names=feature_names, show=)\n\n\ntop_feat = feature_names[np.argsort(np.(sv).mean())[-]]\nshap.dependence_plot(top_feat, sv, X_test_trans, feature_names=feature_names)\n</code></pre>\n<h1>8) Results and Conclusions (example text)</h1>\n<pre><code>\n\n LightGBM (after tuning) achieved the highest ROC AUC on the holdout set:  (example number). XGBoost performed close behind with , and Decision Tree lagged at .\n Feature importance &amp; SHAP analysis indicates that , , and  (example features) are among the most predictive.\n Confusion matrix shows a tendency to predict the majority class — consider class weighting, threshold calibration, or using probability outputs for betting/odds.\n\n\n Tree ensembles (LightGBM/XGBoost) outperform a single Decision Tree.\n Creating difference features between opponents is often powerful for head-to-head prediction.\n Use SHAP to validate that model patterns make sense (domain knowledge: strikes, takedowns, accuracy, cardio proxies matter).\n</code></pre>\n<h1>9) Next steps and Caveats</h1>\n<pre><code>\n Add temporal features: momentum (last N fights), rest days since last fight, age and reach trends.\n Use fighter embedding: train a representation of each fighter using their historical fight sequences.\n Incorporate bookmaker odds as a strong baseline feature.\n Calibrate probabilities using isotonic or Platt scaling if probabilities will be used for decision-making.\n Cross-event leakage: be careful when features leak future info (e.g. post-fight stats).\n\n\n Quality of predictions depends heavily on dataset completeness and feature engineering.\n Avoid leaking future information into training features; always compute features using only data available before the fight.\n</code></pre>",
      "rawMarkdown": "# UFC Fight Prediction Analysis\n\n\n## Table of Contents\n\n1. Required Libraries\n2. Data Loading and Preprocessing\n3. Exploratory Data Analysis\n4. Feature Engineering\n5. Model Implementation (Decision Tree, XGBoost, LightGBM)\n6. Model Comparison and Evaluation\n7. Advanced Visualization and Analysis (SHAP, Lift/Calibration, Partial Dependence)\n8. Results and Conclusions\n9. Next steps and Caveats\n\n---\n\n> **Notes:**\n>\n> * This document is a drop-in Python script / notebook. Replace `data_path` with your dataset path. The code assumes a tabular fight-level dataset where each row is a fight and available columns include fighter-level stats (e.g. `red_*`, `blue_*` or `fighter1_*`/`fighter2_*`), and an outcome column like `winner` or `result`. Adjust column names as needed.\n\n---\n\n# 1) Required Libraries\n\n```python\n# Data + utils\nimport os\nimport numpy as np\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split, StratifiedKFold, cross_val_score, GridSearchCV\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\n\n# Models\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score, confusion_matrix, classification_report, log_loss\nimport xgboost as xgb\nimport lightgbm as lgb\n\n# Explainability & visualization\nimport shap\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# For reproducibility\nRANDOM_STATE = 42\n\n# Optional: silence warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n```\n\n# 2) Data Loading and Preprocessing\n\n```python\n# Load data - update path\ndata_path = 'ufc_fights.csv'  # <- replace\ndf = pd.read_csv(data_path)\n\nprint(df.shape)\nprint(df.columns.tolist()[:60])\n\n# Example target transformation: create binary label `red_wins` if `winner` column lists 'Red'\nif 'winner' in df.columns:\n    df['target'] = (df['winner'].str.lower().str.contains('red')).astype(int)\nelif 'result' in df.columns:\n    # customize based on dataset\n    df['target'] = df['result'].map({'Red':1, 'Blue':0})\nelse:\n    raise ValueError('No known outcome column found. Provide `winner` or `result`.')\n\n# Quick null checks\nprint('Missing per column (top 20):')\nprint(df.isna().sum().sort_values(ascending=False).head(20))\n\n# Drop columns that are identifiers or text not useful for modeling (customize)\n# e.g. fighter names, event text\nto_drop = ['event', 'date', 'fighter_red_name', 'fighter_blue_name']\nfor c in to_drop:\n    if c in df.columns:\n        df.drop(columns=c, inplace=True)\n\n# Simple imputation example: numeric -> fillna median; categorical -> fillna 'missing'\nnum_cols = df.select_dtypes(include=['int64','float64']).columns.tolist()\n# exclude target\nnum_cols = [c for c in num_cols if c != 'target']\ncat_cols = df.select_dtypes(include=['object', 'category']).columns.tolist()\n\nprint('Numeric cols:', num_cols[:10])\nprint('Categorical cols:', cat_cols[:10])\n\n# Basic imputation\nfor c in num_cols:\n    df[c] = df[c].fillna(df[c].median())\nfor c in cat_cols:\n    df[c] = df[c].fillna('missing')\n\n# Train/test split (stratify by target)\ntrain_df, test_df = train_test_split(df, test_size=0.2, stratify=df['target'], random_state=RANDOM_STATE)\n\nX_train = train_df.drop(columns=['target'])\ny_train = train_df['target']\nX_test = test_df.drop(columns=['target'])\ny_test = test_df['target']\n\nprint('Train/test sizes:', X_train.shape, X_test.shape)\n```\n\n# 3) Exploratory Data Analysis (brief)\n\n```python\n# Distribution of target\nprint(y_train.value_counts(normalize=True))\n\n# Correlation heatmap for numeric features (subset)\nplt.figure(figsize=(12,8))\nnum_subset = num_cols[:30]  # limit for plotting\nsns.heatmap(train_df[num_subset + ['target']].corr(), annot=False, cmap='coolwarm')\nplt.title('Numeric feature correlations')\nplt.show()\n\n# Boxplots of some impact features vs target\nfor c in num_cols[:6]:\n    plt.figure(figsize=(6,3))\n    sns.boxplot(x=train_df['target'], y=train_df[c])\n    plt.title(f'{c} by target')\n    plt.show()\n\n# Categorical cardinality\nfor c in cat_cols:\n    print(c, train_df[c].nunique())\n    if train_df[c].nunique() < 20:\n        display(train_df.groupby(c)['target'].mean().sort_values())\n```\n\n# 4) Feature Engineering\n\n```python\n# Example: create differences between fighter stats (red - blue)\n# This depends on your column naming conventions. Example: 'red_sig_strikes', 'blue_sig_strikes'\npairs = []\nfor col in df.columns:\n    if col.startswith('red_'):\n        counterpart = 'blue_' + col[len('red_'):]\n        if counterpart in df.columns:\n            pairs.append((col, counterpart))\n\nfor r,b in pairs:\n    newcol = r.replace('red_','diff_')\n    df[newcol] = df[r] - df[b]\n\n# Recompute train/test after engineering\ntrain_df, test_df = train_test_split(df, test_size=0.2, stratify=df['target'], random_state=RANDOM_STATE)\nX_train = train_df.drop(columns=['target'])\ny_train = train_df['target']\nX_test = test_df.drop(columns=['target'])\ny_test = test_df['target']\n\n# Define final feature lists\nall_num = X_train.select_dtypes(include=['int64','float64']).columns.tolist()\nall_cat = X_train.select_dtypes(include=['object','category']).columns.tolist()\nprint('Final numeric features:', len(all_num))\nprint('Final categorical features:', len(all_cat))\n```\n\n# 5) Model Implementation\n\nWe'll build three pipelines: Decision Tree (sklearn), XGBoost, LightGBM. We'll use a ColumnTransformer to scale numeric features and one-hot encode low-cardinality categoricals.\n\n```python\n# Preprocessing pipeline\nnumeric_transformer = Pipeline(steps=[('scaler', StandardScaler())])\n\n# One-hot encode categoricals with low cardinality, otherwise drop or frequency encode (simple approach)\nlow_cardinality = [c for c in all_cat if X_train[c].nunique() < 20]\nhigh_cardinality = [c for c in all_cat if X_train[c].nunique() >= 20]\n\ncategorical_transformer = Pipeline(steps=[('onehot', OneHotEncoder(handle_unknown='ignore', sparse=False))])\n\npreprocessor = ColumnTransformer(transformers=[\n    ('num', numeric_transformer, all_num),\n    ('cat', categorical_transformer, low_cardinality)\n], remainder='drop')\n\n# Helper: evaluate model function\nfrom sklearn.model_selection import cross_validate\n\ndef evaluate_model(pipe, X, y, cv=5):\n    scoring = ['accuracy','precision','recall','f1','roc_auc']\n    cv_res = cross_validate(pipe, X, y, cv=cv, scoring=scoring, return_train_score=False)\n    return {k: np.mean(v) for k,v in cv_res.items()}\n\n# 5a Decision Tree pipeline\ndt = DecisionTreeClassifier(random_state=RANDOM_STATE)\npipe_dt = Pipeline(steps=[('pre', preprocessor), ('clf', dt)])\n\n# quick baseline\nprint('Decision Tree CV:', evaluate_model(pipe_dt, X_train, y_train, cv=5))\n\n# 5b XGBoost\nxgb_clf = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=RANDOM_STATE)\npipe_xgb = Pipeline(steps=[('pre', preprocessor), ('clf', xgb_clf)])\nprint('XGBoost CV:', evaluate_model(pipe_xgb, X_train, y_train, cv=5))\n\n# 5c LightGBM\nlgb_clf = lgb.LGBMClassifier(random_state=RANDOM_STATE)\npipe_lgb = Pipeline(steps=[('pre', preprocessor), ('clf', lgb_clf)])\nprint('LightGBM CV:', evaluate_model(pipe_lgb, X_train, y_train, cv=5))\n```\n\n## Hyperparameter tuning (example grids)\n\n```python\n# Small grids for demonstration\ndt_param_grid = {\n    'clf__max_depth': [3,6,10,None],\n    'clf__min_samples_leaf': [1,5,10]\n}\n\nxgb_param_grid = {\n    'clf__n_estimators': [100,300],\n    'clf__max_depth': [3,6],\n    'clf__learning_rate': [0.01, 0.1]\n}\n\nlgb_param_grid = {\n    'clf__n_estimators': [100,300],\n    'clf__num_leaves': [31, 63],\n    'clf__learning_rate': [0.01, 0.1]\n}\n\n# Example: GridSearch for LightGBM (smaller cv to save time)\ngs_lgb = GridSearchCV(pipe_lgb, lgb_param_grid, cv=3, scoring='roc_auc', n_jobs=-1, verbose=1)\ngs_lgb.fit(X_train, y_train)\nprint('Best LGB params:', gs_lgb.best_params_)\nprint('Best CV score:', gs_lgb.best_score_)\n\n# Save best estimator\nbest_lgb = gs_lgb.best_estimator_\n\n# Optionally do the same for xgboost and decision tree\n```\n\n# 6) Model Comparison and Evaluation (holdout test)\n\n```python\nmodels = {\n    'DecisionTree': pipe_dt.fit(X_train, y_train),\n    'XGBoost': pipe_xgb.fit(X_train, y_train),\n    'LightGBM': best_lgb if 'best_lgb' in locals() else pipe_lgb.fit(X_train, y_train)\n}\n\nresults = {}\nfor name, model in models.items():\n    y_pred = model.predict(X_test)\n    y_proba = model.predict_proba(X_test)[:,1]\n    results[name] = {\n        'accuracy': accuracy_score(y_test, y_pred),\n        'precision': precision_score(y_test, y_pred),\n        'recall': recall_score(y_test, y_pred),\n        'f1': f1_score(y_test, y_pred),\n        'roc_auc': roc_auc_score(y_test, y_proba),\n        'log_loss': log_loss(y_test, y_proba)\n    }\n\npd.DataFrame(results).T.sort_values('roc_auc', ascending=False)\n\n# Confusion matrix for top model\nbest_name = max(results.keys(), key=lambda k: results[k]['roc_auc'])\nprint('Best model by ROC AUC:', best_name)\nbest_model = models[best_name]\n\ny_pred = best_model.predict(X_test)\ncm = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(5,4))\nsns.heatmap(cm, annot=True, fmt='d')\nplt.title(f'Confusion matrix - {best_name}')\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.show()\n```\n\n# 7) Advanced Visualization and Analysis\n\n```python\n# SHAP explanations for tree-based best model\n# We need to extract the preprocessed matrix and model object depending on pipeline\n# Get feature names after preprocessing\npre = best_model.named_steps['pre']\nclf = best_model.named_steps['clf']\n\n# Fit preprocessor separately to get transformed feature names\npre.fit(X_train)\n# numeric feature names\nnum_names = all_num\ncat_encoder = pre.named_transformers_.get('cat')\nif cat_encoder is not None:\n    cat_names = list(cat_encoder.named_steps['onehot'].get_feature_names_out(low_cardinality))\nelse:\n    cat_names = []\nfeature_names = num_names + cat_names\n\n# Transform test set\nX_test_trans = pre.transform(X_test)\n\n# Use shap TreeExplainer for tree models\nexplainer = shap.TreeExplainer(clf)\nshap_values = explainer.shap_values(X_test_trans)\n\n# For binary classification shap_values is list of arrays; pick class 1\nif isinstance(shap_values, list):\n    sv = shap_values[1]\nelse:\n    sv = shap_values\n\n# Summary plot\nshap.summary_plot(sv, features=X_test_trans, feature_names=feature_names, show=True)\n\n# Dependence plot for top feature\ntop_feat = feature_names[np.argsort(np.abs(sv).mean(0))[-1]]\nshap.dependence_plot(top_feat, sv, X_test_trans, feature_names=feature_names)\n```\n\n# 8) Results and Conclusions (example text)\n\n```markdown\n**Results summary (example):**\n\n- LightGBM (after tuning) achieved the highest ROC AUC on the holdout set: **0.78** (example number). XGBoost performed close behind with **0.76**, and Decision Tree lagged at **0.65**.\n- Feature importance & SHAP analysis indicates that `diff_sig_strikes`, `red_significant_strikes`, and `fighter_experience_diff` (example features) are among the most predictive.\n- Confusion matrix shows a tendency to predict the majority class — consider class weighting, threshold calibration, or using probability outputs for betting/odds.\n\n**Key takeaways:**\n- Tree ensembles (LightGBM/XGBoost) outperform a single Decision Tree.\n- Creating difference features between opponents is often powerful for head-to-head prediction.\n- Use SHAP to validate that model patterns make sense (domain knowledge: strikes, takedowns, accuracy, cardio proxies matter).\n```\n\n# 9) Next steps and Caveats\n\n```markdown\n**Next steps:**\n- Add temporal features: momentum (last N fights), rest days since last fight, age and reach trends.\n- Use fighter embedding: train a representation of each fighter using their historical fight sequences.\n- Incorporate bookmaker odds as a strong baseline feature.\n- Calibrate probabilities using isotonic or Platt scaling if probabilities will be used for decision-making.\n- Cross-event leakage: be careful when features leak future info (e.g. post-fight stats).\n\n**Caveats:**\n- Quality of predictions depends heavily on dataset completeness and feature engineering.\n- Avoid leaking future information into training features; always compute features using only data available before the fight.\n\n```\n\n",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3295343,
      "author_name": "Khushi Yadav",
      "author_url": "",
      "post_date": "2025-09-28T13:31:56.957000",
      "content": "<p>Great Work.@Vicens Gaitan</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 3295118,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-09-27T18:09:21.427000",
      "content": "<p>Great Work. I appreciate 🥳</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3294493,
      "author_name": "Oleh Kivernyk",
      "author_url": "",
      "post_date": "2025-09-26T09:17:21.213000",
      "content": "<p>Nice summary!<br>\nI also used a Minuit (migrad + hesse) + Batman combination. The L-BFGS-B underperformed significantly in my case. However, I used the stellar parameters not only for initialization but also to constrain some of the parameters with Gaussian priors. Did you use the stellar information just for initialization, or did you also apply it as constraints?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3294504,
          "author_name": "Vicens Gaitan",
          "author_url": "",
          "post_date": "2025-09-26T09:35:33.377000",
          "content": "<p>Only for initialization,  but it was important to limit the variation range for each of them based on physical considerations.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3294510,
              "author_name": "Oleh Kivernyk",
              "author_url": "",
              "post_date": "2025-09-26T09:47:12.930000",
              "content": "<p>Thanks for the clarification! I asked it because I'm trying to figure out whether the private star info data was the reason my score dropped so drastically :) </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3295339,
              "author_name": "Oleh Kivernyk",
              "author_url": "",
              "post_date": "2025-09-28T13:18:19.703000",
              "content": "<p>Just curious, did you fit the <code>e</code> and <code>w</code> parameters (eccentricity and longitude of periastron), or are their values fixed in global combined fit?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3295404,
              "author_name": "Vicens Gaitan",
              "author_url": "",
              "post_date": "2025-09-28T16:53:25.667000",
              "content": "<p>I only fit the orbital parametres in the combined fit and fix it  for the 282 airs  frequency fits. You can take a look at the code  <a href=\"https://www.kaggle.com/code/vicensgaitan/6th-place-solution-pipeline\" target=\"_blank\">here</a> in he fit2.py module </p>\n<p>`</p>\n<pre><code>def fit_w(air, minuit_result, fixed_list, dipa_mean, =1, =2):\n\n\n\n\n\n    initial_guesses = DEFAULT_PARAMS_w.copy()\n    default_limits = DEFAULT_LIMITS.copy()\n\n    # Use results  the combined fit as initial guesses\n     p  initial_guesses:\n         p   fixed_list:\n            initial_guesses[p] = minuit_result.values[p]\n\n    # Fix all parameters except depth ()  normalization ()\n    params_to_fix = list(initial_guesses.keys())\n    params_to_fix.()\n    params_to_fix.()\n</code></pre>\n<p>`</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3294274,
      "author_name": "particlebbq",
      "author_url": "",
      "post_date": "2025-09-25T18:10:51.700000",
      "content": "<p>Thanks for posting this write-up!  I like the approach, and the speed you're seeing for the wavelength-by-wavelength fit is amazing -- it strikes me as a really impressive achievement, so congratulations on that in particular (as well as the overall success of your analysis of course!).  I was also using a wavelength-by-wavelength fit in my solution, but I struggled a lot with speed (perhaps because I also allowed free parameters for the ingress/egress duration and limb darkening for each wavelength).  </p>\n<p>Mind if I ask a question about your experience setting up this analysis?  One of the challenges I ran into (right at the end, unfortunately, so I didn't manage to fix it) was a drop in the performance of the minimizers I was using when I tried to port my fitting code to the kaggle notebook environment.  It looked like the same call to scipy.optimize.minimize() just…didn't find as good a minimum in the kaggle notebook as it did in my offline code.  Same numpy/scipy/torch versions online and offline, same input data as far as I could see, same arguments to the function invocation, same gradients (at least, as far as I could spot in the first call to the chisquare function), just….different results, as though the fitter just decided to quit early.    I am wondering if you saw anything similar, and if so, how did you resolve it?  I would love to know what to check for next time.  😅</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3294437,
          "author_name": "Vicens Gaitan",
          "author_url": "",
          "post_date": "2025-09-26T07:21:16.383000",
          "content": "<p>Thanks so much for your feedback! 😊</p>\n<p>In my experience, the results from Scipy optimizers (and also in the case of Minuit) are reproducible across different platforms if both the optimizer's parameters and the data types of the function to be optimized and its parameters are kept the same. Although the optimization process is essentially sequential, some of the intermediate calculations might use a parallel backend (for example, OpenBLAS to compute the Hessian). In this specific scenario, we can't guarantee the exact same result, though it shouldn't differ too significantly. Perhaps this isn't the root of the issue, however.</p>\n<p>In my specific case, using Batman to model the transit, it's been critical to impose physical limits on the orbital and limb darkening values. This prevents the optimizer from getting lost and attempting to use the degrees of freedom from the polynomial drift to explain phenomena that actually depend on the orbital parameters.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3295521,
              "author_name": "particlebbq",
              "author_url": "",
              "post_date": "2025-09-29T02:38:01.913000",
              "content": "<p>Thanks!  There are some good leads here -- I do notice that numpy.show_config() reports a different blas backend on my desktop compared to the kaggle environment; I haven't yet tried switching my offline code to the one kaggle uses, but it's possible that's what's going on here.  I do have limits on the parameters that govern my signal shape, but not on the polynomial background.  I can play around with that a bit, too.</p>\n<p>But I'll definitely have to give iMinuit a shot; I've used the version in ROOT a lot in the past, but I switched to scipy for this contest because I didn't want to try installing ROOT in a Kaggle notebook.  Glad to know that Minuit is available as a pip install now!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3294174,
      "author_name": "Prince Rajak",
      "author_url": "",
      "post_date": "2025-09-25T14:26:37.560000",
      "content": "<h1>UFC Fight Prediction Analysis</h1>\n<h2>Table of Contents</h2>\n<ol>\n<li>Required Libraries</li>\n<li>Data Loading and Preprocessing</li>\n<li>Exploratory Data Analysis</li>\n<li>Feature Engineering</li>\n<li>Model Implementation (Decision Tree, XGBoost, LightGBM)</li>\n<li>Model Comparison and Evaluation</li>\n<li>Advanced Visualization and Analysis (SHAP, Lift/Calibration, Partial Dependence)</li>\n<li>Results and Conclusions</li>\n<li>Next steps and Caveats</li>\n</ol>\n<hr>\n<blockquote>\n  <p><strong>Notes:</strong></p>\n  <ul>\n  <li>This document is a drop-in Python script / notebook. Replace <code>data_path</code> with your dataset path. The code assumes a tabular fight-level dataset where each row is a fight and available columns include fighter-level stats (e.g. <code>red_*</code>, <code>blue_*</code> or <code>fighter1_*</code>/<code>fighter2_*</code>), and an outcome column like <code>winner</code> or <code>result</code>. Adjust column names as needed.</li>\n  </ul>\n</blockquote>\n<hr>\n<h1>1) Required Libraries</h1>\n<pre><code>\n os\n numpy  np\n pandas  pd\n sklearn.model_selection  train_test_split, StratifiedKFold, cross_val_score, GridSearchCV\n sklearn.preprocessing  StandardScaler, OneHotEncoder\n sklearn.compose  ColumnTransformer\n sklearn.pipeline  Pipeline\n\n\n sklearn.tree  DecisionTreeClassifier\n sklearn.metrics  accuracy_score, precision_score, recall_score, f1_score, roc_auc_score, confusion_matrix, classification_report, log_loss\n xgboost  xgb\n lightgbm  lgb\n\n\n shap\n matplotlib.pyplot  plt\n seaborn  sns\n\n\nRANDOM_STATE = \n\n\n warnings\nwarnings.filterwarnings()\n</code></pre>\n<h1>2) Data Loading and Preprocessing</h1>\n<pre><code>\ndata_path =   \ndf = pd.read_csv(data_path)\n\n(df.shape)\n(df.columns.tolist()[:])\n\n\n   df.columns:\n    df[] = (df[]..lower()..contains()).astype()\n   df.columns:\n    \n    df[] = df[].({:, :})\n:\n     ValueError()\n\n\n()\n(df.isna().().sort_values(ascending=).head())\n\n\n\nto_drop = [, , , ]\n c  to_drop:\n     c  df.columns:\n        df.drop(columns=c, inplace=)\n\n\nnum_cols = df.select_dtypes(include=[,]).columns.tolist()\n\nnum_cols = [c  c  num_cols  c != ]\ncat_cols = df.select_dtypes(include=[, ]).columns.tolist()\n\n(, num_cols[:])\n(, cat_cols[:])\n\n\n c  num_cols:\n    df[c] = df[c].fillna(df[c].median())\n c  cat_cols:\n    df[c] = df[c].fillna()\n\n\ntrain_df, test_df = train_test_split(df, test_size=, stratify=df[], random_state=RANDOM_STATE)\n\nX_train = train_df.drop(columns=[])\ny_train = train_df[]\nX_test = test_df.drop(columns=[])\ny_test = test_df[]\n\n(, X_train.shape, X_test.shape)\n</code></pre>\n<h1>3) Exploratory Data Analysis (brief)</h1>\n<pre><code>\n(y_train.value_counts(normalize=))\n\n\nplt.figure(figsize=(,))\nnum_subset = num_cols[:]  \nsns.heatmap(train_df[num_subset + []].corr(), annot=, cmap=)\nplt.title()\nplt.show()\n\n\n c  num_cols[:]:\n    plt.figure(figsize=(,))\n    sns.boxplot(x=train_df[], y=train_df[c])\n    plt.title()\n    plt.show()\n\n\n c  cat_cols:\n    (c, train_df[c].nunique())\n     train_df[c].nunique() &lt; :\n        display(train_df.groupby(c)[].mean().sort_values())\n</code></pre>\n<h1>4) Feature Engineering</h1>\n<pre><code>\n\npairs = []\n col  df.columns:\n     col.startswith():\n        counterpart =  + col[():]\n         counterpart  df.columns:\n            pairs.append((col, counterpart))\n\n r,b  pairs:\n    newcol = r.replace(,)\n    df[newcol] = df[r] - df[b]\n\n\ntrain_df, test_df = train_test_split(df, test_size=, stratify=df[], random_state=RANDOM_STATE)\nX_train = train_df.drop(columns=[])\ny_train = train_df[]\nX_test = test_df.drop(columns=[])\ny_test = test_df[]\n\n\nall_num = X_train.select_dtypes(include=[,]).columns.tolist()\nall_cat = X_train.select_dtypes(include=[,]).columns.tolist()\n(, (all_num))\n(, (all_cat))\n</code></pre>\n<h1>5) Model Implementation</h1>\n<p>We'll build three pipelines: Decision Tree (sklearn), XGBoost, LightGBM. We'll use a ColumnTransformer to scale numeric features and one-hot encode low-cardinality categoricals.</p>\n<pre><code>\nnumeric_transformer = Pipeline(steps=[(, StandardScaler())])\n\n\nlow_cardinality = [c  c  all_cat  X_train[c].nunique() &lt; ]\nhigh_cardinality = [c  c  all_cat  X_train[c].nunique() &gt;= ]\n\ncategorical_transformer = Pipeline(steps=[(, OneHotEncoder(handle_unknown=, sparse=))])\n\npreprocessor = ColumnTransformer(transformers=[\n    (, numeric_transformer, all_num),\n    (, categorical_transformer, low_cardinality)\n], remainder=)\n\n\n sklearn.model_selection  cross_validate\n\n ():\n    scoring = [,,,,]\n    cv_res = cross_validate(pipe, X, y, cv=cv, scoring=scoring, return_train_score=)\n     {k: np.mean(v)  k,v  cv_res.items()}\n\n\ndt = DecisionTreeClassifier(random_state=RANDOM_STATE)\npipe_dt = Pipeline(steps=[(, preprocessor), (, dt)])\n\n\n(, evaluate_model(pipe_dt, X_train, y_train, cv=))\n\n\nxgb_clf = xgb.XGBClassifier(use_label_encoder=, eval_metric=, random_state=RANDOM_STATE)\npipe_xgb = Pipeline(steps=[(, preprocessor), (, xgb_clf)])\n(, evaluate_model(pipe_xgb, X_train, y_train, cv=))\n\n\nlgb_clf = lgb.LGBMClassifier(random_state=RANDOM_STATE)\npipe_lgb = Pipeline(steps=[(, preprocessor), (, lgb_clf)])\n(, evaluate_model(pipe_lgb, X_train, y_train, cv=))\n</code></pre>\n<h2>Hyperparameter tuning (example grids)</h2>\n<pre><code>\ndt_param_grid = {\n    : [,,,],\n    : [,,]\n}\n\nxgb_param_grid = {\n    : [,],\n    : [,],\n    : [, ]\n}\n\nlgb_param_grid = {\n    : [,],\n    : [, ],\n    : [, ]\n}\n\n\ngs_lgb = GridSearchCV(pipe_lgb, lgb_param_grid, cv=, scoring=, n_jobs=-, verbose=)\ngs_lgb.fit(X_train, y_train)\n(, gs_lgb.best_params_)\n(, gs_lgb.best_score_)\n\n\nbest_lgb = gs_lgb.best_estimator_\n\n\n</code></pre>\n<h1>6) Model Comparison and Evaluation (holdout test)</h1>\n<pre><code>models = {\n    : pipe_dt.fit(X_train, y_train),\n    : pipe_xgb.fit(X_train, y_train),\n    : best_lgb    ()  pipe_lgb.fit(X_train, y_train)\n}\n\nresults = {}\n name, model  models.items():\n    y_pred = model.predict(X_test)\n    y_proba = model.predict_proba(X_test)[:,]\n    results[name] = {\n        : accuracy_score(y_test, y_pred),\n        : precision_score(y_test, y_pred),\n        : recall_score(y_test, y_pred),\n        : f1_score(y_test, y_pred),\n        : roc_auc_score(y_test, y_proba),\n        : log_loss(y_test, y_proba)\n    }\n\npd.DataFrame(results).T.sort_values(, ascending=)\n\n\nbest_name = (results.keys(), key= k: results[k][])\n(, best_name)\nbest_model = models[best_name]\n\ny_pred = best_model.predict(X_test)\ncm = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(,))\nsns.heatmap(cm, annot=, fmt=)\nplt.title()\nplt.xlabel()\nplt.ylabel()\nplt.show()\n</code></pre>\n<h1>7) Advanced Visualization and Analysis</h1>\n<pre><code>\n\n\npre = best_model.named_steps[]\nclf = best_model.named_steps[]\n\n\npre.fit(X_train)\n\nnum_names = all_num\ncat_encoder = pre.named_transformers_.get()\n cat_encoder   :\n    cat_names = (cat_encoder.named_steps[].get_feature_names_out(low_cardinality))\n:\n    cat_names = []\nfeature_names = num_names + cat_names\n\n\nX_test_trans = pre.transform(X_test)\n\n\nexplainer = shap.TreeExplainer(clf)\nshap_values = explainer.shap_values(X_test_trans)\n\n\n (shap_values, ):\n    sv = shap_values[]\n:\n    sv = shap_values\n\n\nshap.summary_plot(sv, features=X_test_trans, feature_names=feature_names, show=)\n\n\ntop_feat = feature_names[np.argsort(np.(sv).mean())[-]]\nshap.dependence_plot(top_feat, sv, X_test_trans, feature_names=feature_names)\n</code></pre>\n<h1>8) Results and Conclusions (example text)</h1>\n<pre><code>\n\n LightGBM (after tuning) achieved the highest ROC AUC on the holdout set:  (example number). XGBoost performed close behind with , and Decision Tree lagged at .\n Feature importance &amp; SHAP analysis indicates that , , and  (example features) are among the most predictive.\n Confusion matrix shows a tendency to predict the majority class — consider class weighting, threshold calibration, or using probability outputs for betting/odds.\n\n\n Tree ensembles (LightGBM/XGBoost) outperform a single Decision Tree.\n Creating difference features between opponents is often powerful for head-to-head prediction.\n Use SHAP to validate that model patterns make sense (domain knowledge: strikes, takedowns, accuracy, cardio proxies matter).\n</code></pre>\n<h1>9) Next steps and Caveats</h1>\n<pre><code>\n Add temporal features: momentum (last N fights), rest days since last fight, age and reach trends.\n Use fighter embedding: train a representation of each fighter using their historical fight sequences.\n Incorporate bookmaker odds as a strong baseline feature.\n Calibrate probabilities using isotonic or Platt scaling if probabilities will be used for decision-making.\n Cross-event leakage: be careful when features leak future info (e.g. post-fight stats).\n\n\n Quality of predictions depends heavily on dataset completeness and feature engineering.\n Avoid leaking future information into training features; always compute features using only data available before the fight.\n</code></pre>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3294167": "# Ariel 25 Challenge: A Three-Stage Pipeline for Robust Transmission Spectra\n\nI'd first like to thank the organizers for putting on such an interesting competition. They deserve credit both for their scientific rigor and for the intricate puzzles they used to generate the data—some of which I still haven't managed to solve.\n\n##1. Summary\n\nThis solution presents a comprehensive, three-stage pipeline designed to extract high-fidelity transmission spectra from the FGS1 and AIR-CH0 Ariel instrument data. The approach combines physics-based modeling with a machine learning-inspired post-processing framework to tackle the key challenges of instrumental systematics and uncertainty estimation.\n\n1. Stage 1: GPU-Accelerated Preprocessing. Raw sensor data is converted into clean, calibrated light curves using a CuPy-based pipeline. This stage handles standard instrumental corrections, including non-linearity, dark current, flat-fielding, and jitter, all performed efficiently on the GPU.\n\n2. Stage 2: Hierarchical Transit Fitting. I use the [**batman**](https://lkreidberg.github.io/batman/docs/html/index.html) transit modeling library and the [**iminuit**](https://scikit-hep.org/iminuit/) optimizer to perform a multi-step fit. A robust global model is first fit to the combined FGS and AIRS light curves to constrain key orbital parameters. This is followed by a 2D drift correction and a wavelength-by-wavelength fit to extract the initial transmission spectrum.\n\n3. Stage 3: Cross-Validated Post-Processing Ensemble. The initial spectra are refined using an ensemble of models trained with Grouped K-Fold Cross-Validation. This final stage blends a PCA-regularized signal with a smoothed signal and employs a parameterized model to predict the final uncertainties (sigmas), optimizing all parameters directly against the competition score.\n\n\nThis multi-stage design ensures that physical constraints are respected in the initial fit, while the final model has the flexibility to learn and correct for residual systematic errors and produce well-calibrated uncertainties.\n\n--------------------------------------------------------------------------------\n## 2. Methodology\n\n###Stage 1: Preprocessing Raw Data\nThe first step in the pipeline is to transform the raw sensor readouts into scientifically useful light curves. This entire process is executed on the GPU using CuPy for maximum efficiency.\n\nKey Preprocessing Steps:\n\n• Initial Calibrations: The pipeline begins by applying standard instrumental corrections. This includes an ADC correction for gain and offset, a 5th-degree polynomial correction for detector non-linearity (apply_linear_corr_fast), and subtraction of a scaled master dark frame (clean_dark).\n\n• Flat-Fielding & Masking: A master flat frame is applied to correct for pixel-to-pixel sensitivity variations. Pixels identified as \"hot\" (via sigma-clipping the dark frame) or \"dead\" are masked.\n\n• Jitter Correction: For the FGS1 sensor, we apply a center-of-mass  regression to correct for flux variations caused by image motion on the detector. The flux is de-correlated from the normalized X and Y  positions of the stellar image. While a PCA-based jitter correction method was developed for the AIRS sensor, it was found to be less effective and is disabled in the final pipeline.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2Fc209358dd02c309ac5067f60b7903724%2Fjitter.png?generation=1758796485996693&alt=media)\n\n\n\n• Signal Extraction & Cleaning:\n\n   \n* For FGS1, the final flux is extracted using simple aperture photometry. PSF photometry gives lower SNR signals, probably due to the jitter.\n* For AIRS, the signal is extracted from the central detector region, and a background signal, calculated from the top and bottom edges of the detector, is subtracted. This is not optimal because the spectroscopy signal extends i a difractive way to the whole sensor, making dificult to  identify the aditive frequency  dependent background and  making necessary a global  scale factor and an spectrum slope correction in the post processing stage\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F7f6d6424c53a3475488d53772ec8b701%2Fair_psf.png?generation=1758806615604977&alt=media)\n* A spike-cleaning algorithm  is applied to the final time series to remove cosmic ray hits by identifying and replacing outliers in the frame-to-frame difference.\n\n    \n• Binning: To improve the signal-to-noise ratio, the final time series for both instruments are binned by averaging consecutive frames.\n\n###Stage 2: Physics-Based Spectral Fitting\n\nWith clean light curves for both instruments, we extract the initial transmission spectrum. This process is parallelized using pqdm to efficiently process all observations.\n\nHierarchical Fitting Strategy:\n\n**1. Global Combined Fit**: We first fit the FGS light curve and the wavelength-averaged AIRS light curve **simultaneously** using fit_combined_curves. This crucial step provides robust constraints on the shared physical parameters of the system (e.g., orbital period per, inclination inc, transit time t0), which are initialized from the provided star_info.csv file.  The transit is modeled using batman, asuming a **quadratic limb darkening model**, using 2 parameters, and instrumental trends are modeled with a polynomial baseline in time.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F04530cea65be4483a14e6effec93bbe2%2Ffits.png?generation=1758796534186505&alt=media)\n\n**2. Refined Fit & Drift Correction**: The results from the first fit are used to initialize a second, refined fit on normalized data. Subsequently, a 2D instrumental drift model is fit to the out-of-transit portion of the AIRS data cube. This model (Drift class) consists of two 4th-order polynomials, one for the frequency and one for the temporal axis, and it corrects for slow-varying systematic patterns across the detector. The data cube is then divided by this drift model.\n**3. Wavelength-by-Wavelength Fit**: The final step is to measure the transit depth in each individual wavelength channel of the drift-corrected AIRS data. The fit_w function iterates through the channels, fitting the data within a small moving window (deltaw=2) to boost signal-to-noise. In this fit, most orbital parameters are fixed to the values from the global fit, and only the transit depth (dipa) and baseline normalization (A0a) are allowed to vary.\nThis process results in the \"raw\" spectrum  and its associated uncertainties , which serve as the primary inputs for our final modeling stage.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F254638%2F990115e5eb8fa6e302f721f552a1cb94%2Fspectrum.png?generation=1758796611221256&alt=media)\n\n## Stage 3: Post-Processing and Uncertainty Modeling\n\nThe final stage of the solution is a post-processing model that refines the raw spectra and, critically, learns a robust model for the final uncertainties. This is where the bulk of the \"machine learning\" occurs.\n\nModel Architecture (build_preds): The model generates a final prediction by blending two different representations of the raw spectrum (preds1):\n\n**PCA-Regularized Prediction (preds0)**: The raw spectrum undergoes a slope correction and is then projected onto a pre-defined PCA basis (components.npy). This de-noises the spectrum and imposes a strong regularization prior, capturing the most common modes of variation.\n**Smoothed Prediction (preds2)**: The raw AIRS spectrum is smoothed using a Savitzky-Golay filter to reduce high-frequency noise while preserving broader spectral features.\n\nThe final mean prediction is a weighted average of these two models: preds = app0 + (1-ap)preds2, where ap is a learned parameter.\n\n**Sigma Model**: Predicting accurate uncertainties is key to maximizing the competition score. I developed a complex, empirically-derived formula to model the final sigma for each data point. This formula combines multiple sources of uncertainty:\n\n \n* a1* err^e1 : The propagated error from the initial fit, with a learned exponent.\n* a01/p1^e2  : A term related to the signal magnitude.\n* b0 * sigma0: A term proportional to the raw signal's volatility.\n* ae * np.abs(p0-preds2): A term that increases uncertainty where the PCA and smoothed models disagree, capturing model \nuncertainty or unknow spectra.\n\nFinally ther is a sigma  adjustment  samples with very few or very many out-of-transit points, or very differnt dip por FGS1 and AIR-CH0\n\n#Training and Ensembling:\n\n  **Objective**: The free parameters of the model (ap, e1, e2, a1, a0, b0, ae, etc.) are optimized using iminuit to directly maximize the competition score.\n\n  **Cross-Validation**: To build a robust model that generalizes well, we employ a 10-fold Grouped K-Fold Cross-Validation strategy, using planet_id to group the data. This ensures that all observations of a single planet are kept within the same fold, preventing data leakage.\n\n  **Calibration**: Within each fold, after the primary parameters are optimized, we calculate a wavelength-dependent calibration factor (alpha) that scales the predicted sigmas to best match the variance observed in the training data. This factor is then applied to the validation set predictions.\n  \n   **Ensembling**: The final submission is generated by averaging the calibrated predictions and sigmas from the 10 models trained during the cross-validation process. This ensembling technique reduces variance and improves the final score.\n\nThe entire training process, including the parallel optimization of each fold, is managed by the CV_model function in the model.py file.\n\n--------------------------------------------------------------------------------\n4. What Did Not Work\nI attempted to post-process the spectrum and sigma values using Gaussian Processes and Ridge Regression, but the results were worse than using the blending of a smoothed/PCA-regularized model and the empirical sigma formula.\n\nIn the final weeks, I spent time trying to isolate the additive foreground noise in the AIRS data for each sample, which would have eliminated the need for the global factor and slope compensation. However, I was unable to find a way to de-correlate signal and noise. I observed a clear structure in the spectroscopic PSF, but all attempts to extract the noise yielded incorrect values that failed to compensate for the systematics.\n\n5. A Comment on Efficiency\nThe choice of batman and iMinuit was driven by speed. iMinuit is a Python wrapper for the C++ port of the MINUIT Fortran code (developed by Fred James in the '70s; [Paper](https://www.sciencedirect.com/science/article/abs/pii/0010465575900399?via%3Dihub), [Wikipedia](https://en.wikipedia.org/wiki/MINUIT)), and it is typically faster than generic SciPy optimization routines. Additionally, the use of GPU acceleration for data processing via CuPy results in a very fast pipeline because all processes are fully parallelized. The submission process on Kaggle using a P100 instance takes less than 1.5 hours, and on a machine with 100 cores and 4  nVidia L40s GPUs, the full preprocess, fitting, and post-process takes around 5 minutes. The post-process parameter adjustment can also be done in minutes.\n\n--------------------------------------------------------------------------------\n##6. Conclusion\nThis solution demonstrates the power of a hybrid approach, blending physics-informed transit modeling with a flexible, data-driven post-processing framework. By first extracting a reasonable physical spectrum and then refining it with a model trained to optimize the specific competition metric, we can effectively correct for complex instrumental systematics and produce highly accurate, well-calibrated transmission spectra. The use of GPU acceleration, parallel processing, and a robust cross-validation scheme ensures that this complex pipeline remains computationally efficient and resistant to overfitting.",
    "3295343": "Great Work.@Vicens Gaitan",
    "3295118": "Great Work. I appreciate 🥳",
    "3294493": "Nice summary!\nI also used a Minuit (migrad + hesse) + Batman combination. The L-BFGS-B underperformed significantly in my case. However, I used the stellar parameters not only for initialization but also to constrain some of the parameters with Gaussian priors. Did you use the stellar information just for initialization, or did you also apply it as constraints?",
    "3294274": "Thanks for posting this write-up!  I like the approach, and the speed you're seeing for the wavelength-by-wavelength fit is amazing -- it strikes me as a really impressive achievement, so congratulations on that in particular (as well as the overall success of your analysis of course!).  I was also using a wavelength-by-wavelength fit in my solution, but I struggled a lot with speed (perhaps because I also allowed free parameters for the ingress/egress duration and limb darkening for each wavelength).  \n\nMind if I ask a question about your experience setting up this analysis?  One of the challenges I ran into (right at the end, unfortunately, so I didn't manage to fix it) was a drop in the performance of the minimizers I was using when I tried to port my fitting code to the kaggle notebook environment.  It looked like the same call to scipy.optimize.minimize() just...didn't find as good a minimum in the kaggle notebook as it did in my offline code.  Same numpy/scipy/torch versions online and offline, same input data as far as I could see, same arguments to the function invocation, same gradients (at least, as far as I could spot in the first call to the chisquare function), just....different results, as though the fitter just decided to quit early.    I am wondering if you saw anything similar, and if so, how did you resolve it?  I would love to know what to check for next time.  😅",
    "3294174": "# UFC Fight Prediction Analysis\n\n\n## Table of Contents\n\n1. Required Libraries\n2. Data Loading and Preprocessing\n3. Exploratory Data Analysis\n4. Feature Engineering\n5. Model Implementation (Decision Tree, XGBoost, LightGBM)\n6. Model Comparison and Evaluation\n7. Advanced Visualization and Analysis (SHAP, Lift/Calibration, Partial Dependence)\n8. Results and Conclusions\n9. Next steps and Caveats\n\n---\n\n> **Notes:**\n>\n> * This document is a drop-in Python script / notebook. Replace `data_path` with your dataset path. The code assumes a tabular fight-level dataset where each row is a fight and available columns include fighter-level stats (e.g. `red_*`, `blue_*` or `fighter1_*`/`fighter2_*`), and an outcome column like `winner` or `result`. Adjust column names as needed.\n\n---\n\n# 1) Required Libraries\n\n```python\n# Data + utils\nimport os\nimport numpy as np\nimport pandas as pd\nfrom sklearn.model_selection import train_test_split, StratifiedKFold, cross_val_score, GridSearchCV\nfrom sklearn.preprocessing import StandardScaler, OneHotEncoder\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\n\n# Models\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score, confusion_matrix, classification_report, log_loss\nimport xgboost as xgb\nimport lightgbm as lgb\n\n# Explainability & visualization\nimport shap\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# For reproducibility\nRANDOM_STATE = 42\n\n# Optional: silence warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n```\n\n# 2) Data Loading and Preprocessing\n\n```python\n# Load data - update path\ndata_path = 'ufc_fights.csv'  # <- replace\ndf = pd.read_csv(data_path)\n\nprint(df.shape)\nprint(df.columns.tolist()[:60])\n\n# Example target transformation: create binary label `red_wins` if `winner` column lists 'Red'\nif 'winner' in df.columns:\n    df['target'] = (df['winner'].str.lower().str.contains('red')).astype(int)\nelif 'result' in df.columns:\n    # customize based on dataset\n    df['target'] = df['result'].map({'Red':1, 'Blue':0})\nelse:\n    raise ValueError('No known outcome column found. Provide `winner` or `result`.')\n\n# Quick null checks\nprint('Missing per column (top 20):')\nprint(df.isna().sum().sort_values(ascending=False).head(20))\n\n# Drop columns that are identifiers or text not useful for modeling (customize)\n# e.g. fighter names, event text\nto_drop = ['event', 'date', 'fighter_red_name', 'fighter_blue_name']\nfor c in to_drop:\n    if c in df.columns:\n        df.drop(columns=c, inplace=True)\n\n# Simple imputation example: numeric -> fillna median; categorical -> fillna 'missing'\nnum_cols = df.select_dtypes(include=['int64','float64']).columns.tolist()\n# exclude target\nnum_cols = [c for c in num_cols if c != 'target']\ncat_cols = df.select_dtypes(include=['object', 'category']).columns.tolist()\n\nprint('Numeric cols:', num_cols[:10])\nprint('Categorical cols:', cat_cols[:10])\n\n# Basic imputation\nfor c in num_cols:\n    df[c] = df[c].fillna(df[c].median())\nfor c in cat_cols:\n    df[c] = df[c].fillna('missing')\n\n# Train/test split (stratify by target)\ntrain_df, test_df = train_test_split(df, test_size=0.2, stratify=df['target'], random_state=RANDOM_STATE)\n\nX_train = train_df.drop(columns=['target'])\ny_train = train_df['target']\nX_test = test_df.drop(columns=['target'])\ny_test = test_df['target']\n\nprint('Train/test sizes:', X_train.shape, X_test.shape)\n```\n\n# 3) Exploratory Data Analysis (brief)\n\n```python\n# Distribution of target\nprint(y_train.value_counts(normalize=True))\n\n# Correlation heatmap for numeric features (subset)\nplt.figure(figsize=(12,8))\nnum_subset = num_cols[:30]  # limit for plotting\nsns.heatmap(train_df[num_subset + ['target']].corr(), annot=False, cmap='coolwarm')\nplt.title('Numeric feature correlations')\nplt.show()\n\n# Boxplots of some impact features vs target\nfor c in num_cols[:6]:\n    plt.figure(figsize=(6,3))\n    sns.boxplot(x=train_df['target'], y=train_df[c])\n    plt.title(f'{c} by target')\n    plt.show()\n\n# Categorical cardinality\nfor c in cat_cols:\n    print(c, train_df[c].nunique())\n    if train_df[c].nunique() < 20:\n        display(train_df.groupby(c)['target'].mean().sort_values())\n```\n\n# 4) Feature Engineering\n\n```python\n# Example: create differences between fighter stats (red - blue)\n# This depends on your column naming conventions. Example: 'red_sig_strikes', 'blue_sig_strikes'\npairs = []\nfor col in df.columns:\n    if col.startswith('red_'):\n        counterpart = 'blue_' + col[len('red_'):]\n        if counterpart in df.columns:\n            pairs.append((col, counterpart))\n\nfor r,b in pairs:\n    newcol = r.replace('red_','diff_')\n    df[newcol] = df[r] - df[b]\n\n# Recompute train/test after engineering\ntrain_df, test_df = train_test_split(df, test_size=0.2, stratify=df['target'], random_state=RANDOM_STATE)\nX_train = train_df.drop(columns=['target'])\ny_train = train_df['target']\nX_test = test_df.drop(columns=['target'])\ny_test = test_df['target']\n\n# Define final feature lists\nall_num = X_train.select_dtypes(include=['int64','float64']).columns.tolist()\nall_cat = X_train.select_dtypes(include=['object','category']).columns.tolist()\nprint('Final numeric features:', len(all_num))\nprint('Final categorical features:', len(all_cat))\n```\n\n# 5) Model Implementation\n\nWe'll build three pipelines: Decision Tree (sklearn), XGBoost, LightGBM. We'll use a ColumnTransformer to scale numeric features and one-hot encode low-cardinality categoricals.\n\n```python\n# Preprocessing pipeline\nnumeric_transformer = Pipeline(steps=[('scaler', StandardScaler())])\n\n# One-hot encode categoricals with low cardinality, otherwise drop or frequency encode (simple approach)\nlow_cardinality = [c for c in all_cat if X_train[c].nunique() < 20]\nhigh_cardinality = [c for c in all_cat if X_train[c].nunique() >= 20]\n\ncategorical_transformer = Pipeline(steps=[('onehot', OneHotEncoder(handle_unknown='ignore', sparse=False))])\n\npreprocessor = ColumnTransformer(transformers=[\n    ('num', numeric_transformer, all_num),\n    ('cat', categorical_transformer, low_cardinality)\n], remainder='drop')\n\n# Helper: evaluate model function\nfrom sklearn.model_selection import cross_validate\n\ndef evaluate_model(pipe, X, y, cv=5):\n    scoring = ['accuracy','precision','recall','f1','roc_auc']\n    cv_res = cross_validate(pipe, X, y, cv=cv, scoring=scoring, return_train_score=False)\n    return {k: np.mean(v) for k,v in cv_res.items()}\n\n# 5a Decision Tree pipeline\ndt = DecisionTreeClassifier(random_state=RANDOM_STATE)\npipe_dt = Pipeline(steps=[('pre', preprocessor), ('clf', dt)])\n\n# quick baseline\nprint('Decision Tree CV:', evaluate_model(pipe_dt, X_train, y_train, cv=5))\n\n# 5b XGBoost\nxgb_clf = xgb.XGBClassifier(use_label_encoder=False, eval_metric='logloss', random_state=RANDOM_STATE)\npipe_xgb = Pipeline(steps=[('pre', preprocessor), ('clf', xgb_clf)])\nprint('XGBoost CV:', evaluate_model(pipe_xgb, X_train, y_train, cv=5))\n\n# 5c LightGBM\nlgb_clf = lgb.LGBMClassifier(random_state=RANDOM_STATE)\npipe_lgb = Pipeline(steps=[('pre', preprocessor), ('clf', lgb_clf)])\nprint('LightGBM CV:', evaluate_model(pipe_lgb, X_train, y_train, cv=5))\n```\n\n## Hyperparameter tuning (example grids)\n\n```python\n# Small grids for demonstration\ndt_param_grid = {\n    'clf__max_depth': [3,6,10,None],\n    'clf__min_samples_leaf': [1,5,10]\n}\n\nxgb_param_grid = {\n    'clf__n_estimators': [100,300],\n    'clf__max_depth': [3,6],\n    'clf__learning_rate': [0.01, 0.1]\n}\n\nlgb_param_grid = {\n    'clf__n_estimators': [100,300],\n    'clf__num_leaves': [31, 63],\n    'clf__learning_rate': [0.01, 0.1]\n}\n\n# Example: GridSearch for LightGBM (smaller cv to save time)\ngs_lgb = GridSearchCV(pipe_lgb, lgb_param_grid, cv=3, scoring='roc_auc', n_jobs=-1, verbose=1)\ngs_lgb.fit(X_train, y_train)\nprint('Best LGB params:', gs_lgb.best_params_)\nprint('Best CV score:', gs_lgb.best_score_)\n\n# Save best estimator\nbest_lgb = gs_lgb.best_estimator_\n\n# Optionally do the same for xgboost and decision tree\n```\n\n# 6) Model Comparison and Evaluation (holdout test)\n\n```python\nmodels = {\n    'DecisionTree': pipe_dt.fit(X_train, y_train),\n    'XGBoost': pipe_xgb.fit(X_train, y_train),\n    'LightGBM': best_lgb if 'best_lgb' in locals() else pipe_lgb.fit(X_train, y_train)\n}\n\nresults = {}\nfor name, model in models.items():\n    y_pred = model.predict(X_test)\n    y_proba = model.predict_proba(X_test)[:,1]\n    results[name] = {\n        'accuracy': accuracy_score(y_test, y_pred),\n        'precision': precision_score(y_test, y_pred),\n        'recall': recall_score(y_test, y_pred),\n        'f1': f1_score(y_test, y_pred),\n        'roc_auc': roc_auc_score(y_test, y_proba),\n        'log_loss': log_loss(y_test, y_proba)\n    }\n\npd.DataFrame(results).T.sort_values('roc_auc', ascending=False)\n\n# Confusion matrix for top model\nbest_name = max(results.keys(), key=lambda k: results[k]['roc_auc'])\nprint('Best model by ROC AUC:', best_name)\nbest_model = models[best_name]\n\ny_pred = best_model.predict(X_test)\ncm = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(5,4))\nsns.heatmap(cm, annot=True, fmt='d')\nplt.title(f'Confusion matrix - {best_name}')\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.show()\n```\n\n# 7) Advanced Visualization and Analysis\n\n```python\n# SHAP explanations for tree-based best model\n# We need to extract the preprocessed matrix and model object depending on pipeline\n# Get feature names after preprocessing\npre = best_model.named_steps['pre']\nclf = best_model.named_steps['clf']\n\n# Fit preprocessor separately to get transformed feature names\npre.fit(X_train)\n# numeric feature names\nnum_names = all_num\ncat_encoder = pre.named_transformers_.get('cat')\nif cat_encoder is not None:\n    cat_names = list(cat_encoder.named_steps['onehot'].get_feature_names_out(low_cardinality))\nelse:\n    cat_names = []\nfeature_names = num_names + cat_names\n\n# Transform test set\nX_test_trans = pre.transform(X_test)\n\n# Use shap TreeExplainer for tree models\nexplainer = shap.TreeExplainer(clf)\nshap_values = explainer.shap_values(X_test_trans)\n\n# For binary classification shap_values is list of arrays; pick class 1\nif isinstance(shap_values, list):\n    sv = shap_values[1]\nelse:\n    sv = shap_values\n\n# Summary plot\nshap.summary_plot(sv, features=X_test_trans, feature_names=feature_names, show=True)\n\n# Dependence plot for top feature\ntop_feat = feature_names[np.argsort(np.abs(sv).mean(0))[-1]]\nshap.dependence_plot(top_feat, sv, X_test_trans, feature_names=feature_names)\n```\n\n# 8) Results and Conclusions (example text)\n\n```markdown\n**Results summary (example):**\n\n- LightGBM (after tuning) achieved the highest ROC AUC on the holdout set: **0.78** (example number). XGBoost performed close behind with **0.76**, and Decision Tree lagged at **0.65**.\n- Feature importance & SHAP analysis indicates that `diff_sig_strikes`, `red_significant_strikes`, and `fighter_experience_diff` (example features) are among the most predictive.\n- Confusion matrix shows a tendency to predict the majority class — consider class weighting, threshold calibration, or using probability outputs for betting/odds.\n\n**Key takeaways:**\n- Tree ensembles (LightGBM/XGBoost) outperform a single Decision Tree.\n- Creating difference features between opponents is often powerful for head-to-head prediction.\n- Use SHAP to validate that model patterns make sense (domain knowledge: strikes, takedowns, accuracy, cardio proxies matter).\n```\n\n# 9) Next steps and Caveats\n\n```markdown\n**Next steps:**\n- Add temporal features: momentum (last N fights), rest days since last fight, age and reach trends.\n- Use fighter embedding: train a representation of each fighter using their historical fight sequences.\n- Incorporate bookmaker odds as a strong baseline feature.\n- Calibrate probabilities using isotonic or Platt scaling if probabilities will be used for decision-making.\n- Cross-event leakage: be careful when features leak future info (e.g. post-fight stats).\n\n**Caveats:**\n- Quality of predictions depends heavily on dataset completeness and feature engineering.\n- Avoid leaking future information into training features; always compute features using only data available before the fight.\n\n```\n\n"
  }
}