{
  "id": 609334,
  "title": "30th place solution",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609334",
  "author_name": "particlebbq",
  "post_date": "2025-09-26T01:04:11.579000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Executive summary</h1>\n<p>A chisquare fit of a parametric function (3- or 5-parameter trapezoid for signal, cubic polynomial for background) to all wavelengths provides inputs to a small, 1D convolutional neural net which smooths the fit results and attempts to correct for limb darkening and any internal fit bias which may be present. The most significant challenge to this approach was getting the fits to achieve adequate convergence within the runtime constraints of the competition, and so this write-up emphasizes timing measurements and time-saving techniques where appropriate.</p>\n<p>A potentially interesting byproduct of this analysis is a wavelength-by-wavelength measurement of ingress/egress time and limb darkening; however, as no ground truth was provided for these quantities, it is difficult to be sure that these measurements, as they are now, are fit for any downstream purpose.</p>\n<h1>Preprocessing</h1>\n<p>The initial processing and calibration of the raw data follows the standard procedure, described in <a href=\"https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data\" target=\"_blank\">this notebook</a>; however, the implementation used in the present solution performs several of these operations on the gpu in order to save time.</p>\n<p>A simple outlier-exclusion procedure is applied to both the FGS and airs data.  Before summing over spatial dimensions, the mean and standard deviation of the readout for each pixel are computed in a sliding window over 50 timesteps. Any time a pixel's readout is more than 3 sigma above or below the mean in that window, its value is replaced by the mean.</p>\n<p>Both the AIRS and FGS data are summed over their respective spatial dimensions to obtain one light curve for each wavelength. FGS light curve are summed in bins of 12 timesteps to match the AIRS light curves, and then both are collected into bins of 15 timesteps, so that each light curve for each detector and wavelength is comprised of 375 time bins.  Each light curve is normalized so that its average over time is 1.  The mean and standard error of the observations in each time bin provide an initial set of y and y uncertainty values for the fit.  In an attempt to improve the convergence of the fit, the y uncertainty values are smoothed by replacing each y uncertainty value with the mean of the uncertainties in a sliding window of 5 time bins.  There was unfortunately not enough time to perform an ablation study to verify the impact of this step on the final version of the analysis before the close of the competition. The x values used in the fit are evenly-spaced numbers ranging from 0 at the start of data collection to 1 at the end.</p>\n<p>An initial detection of the ingress and egress regions is obtained by performing a continuous wavelet transform (using the 'cgau1' wavelet in pywt) of the FGS light curve and searching for maxima in the resulting wavelet coefficients.  The function that does this returns a mean and a width for the two largest peaks in the transformed data.  The initial estimate of the ingress region is taken to be a window of +/- 2 sigma around the earlier of the two peaks, and the intial estimate for the egress region is taken to be a window of +/- 2sigma around the other peak.  The in-transit region is the region between the ingress and egress regions, and the out-of-transit region is the rest of the light curve.  These initial estimates serve two purposes in this analysis:  they are used to compute initial guesses for certain fit parameters, and they are also used in a \"Gaussian baseline\" analysis that is used as a fallback when there are indications that a fit has failed.</p>\n<p>The \"Gaussian baseline\" is a simple cut-and-count estimate of the transit depth.  Given a normalized light curve (from either FGS or from any AIRS wavelength) and the initial estimates of the ingress and egress regions, the baseline analysis computes the fit depth as the difference between the mean light curve value in the in-transit and out-of-transit regions, divided by the mean light curve value in the out-of-transit region.  Gaussian propagation of uncertainty provides a crude estimate of the uncertainty on that fit depth.  This estimate is used later to compute features that are included in the neural net inputs, and it is also used as a replacement for the fit values themselves in a \"fallback\" result that is reported in the submission when the fit results are deemed too unreliable (according to a set of selections defined below).</p>\n<p>The initial estimates of the in-transit and out-of-transit regions are also used to compute initial values for the parameters of the FGS fit described below. The distance between the means is taken as an initial guess for T in the fit; the midpoint is taken as an initial guess for Tcenter; and the average of the two widths is the initial guess for tau.  An initial guess for the limb-darkening parameter is derived from the range of normalized light-curve intensities present in the in-transit region.  The background polynomial parameters are initialized such that the constant term is 1 and the linear, quadratic, and cubic terms are all 0.  The initial guess for the fit depth is taken from the Gaussian baseline.  Initial guesses for the AIRS fits are taken from the output of the FGS fit.</p>\n<p>Timing estimates of the preprocessing phase of the analysis include all of the above operations and are performed in the Kaggle notebook environment with a P100 gpu, usually running on 50-100 examples from the training set.  There was some variability between one timing run and the next, with this phase of the analysis typically clocking in at anywhere between 12 and 15 seconds per transit on average. </p>\n<h1>Parametric fit</h1>\n<p>The signal is modeled as a trapezoid, multiplied by a parabola centered at the trapezoid's midpoint.  Its shape is characterized by five parameters:</p>\n<ul>\n<li>Tcenter, the midpoint of the trapezoid and parabola.</li>\n<li>T, the duration of the transit, i.e. the time between the middle of the ingress (left trapezoid leg) and the middle of the egress (right trapezoid leg)</li>\n<li>tau, the ingress/egress time.  This parameter is shared between the ingress and egress regions; that is, the signal model does not allow for the ingress and egress times to be different.</li>\n<li>LDC, the strength of the parabola which models the effect of limb darkening.  Note that this is not the same as any of the parameters of limb darkening models in the literature; it is more properly understood as an empirical parameter which describes the shape of the top of the trapezoid.</li>\n</ul>\n<p>The values of these parameters in the fit are all limited to fixed ranges by transforming the fit parameters through tanh or sigmoid functions. The trapezoid shape is implemented in a fast, gpu-friendly way using pytorch  (code simplified slightly for readability):</p>\n<p><code>trapezoid=torch.clip((torch.arange(0,n_timesteps)-tstart)/tau,0,1)*torch.clip((tend-torch.arange(0,n_timesteps))/tau,0,1)</code></p>\n<p>where <code>tstart</code> and <code>tend</code> are <code>Tcenter-T/2</code> and <code>Tcenter+T/2</code>, respectively.  The limb-darkening effect is applied as a multiplicative correction to this:</p>\n<pre><code>=torch.es((n_timesteps,))\n=(torch.arange(,n_timesteps)-Tcenter)/max(T/,)\n=torch.clip(limbdark-ldc1*limbdark_x**,,)\n=limbdark*trapezoid\n</code></pre>\n<p>This four-parameter shape is normalized so that its maximum value is 1 and is then combined with the fourth-order polynomial background with the help of another fit parameter, the transit depth:</p>\n<p><code>fit_y=fit_y*(1-fit_depth*signal_shape)</code></p>\n<p>The transit depth is constrained to be positive by transforming the corresponding fit parameter using a softplus function.  Since the model is implemented in pytorch, it is straightforward to provide a gradient estimate to the minimizer using torch's automatic differentiation.  A chisquare function is defined in the usual way as the sum of <code>(obs_y-fit_y)**2/y_err**2</code> over all time and wavelength bins.  The model is first fit to the FGS data using scipy.optimize.minimize() to find the fit parameters that minimize chisquare, allowing all parameters to float freely in the fit.  Up to three calls to minimize() are allowed:  the first uses the default minimizer, BFGS; if that fails (or if the corresponding uncertainty analysis, described below, fails), then a second call is made to the SLSQP minimizer; if that fails, then a third call is made to the Powell minimzer.  Since the FGS data consist of only one wavelength, i.e. only one light curve, this fit does not benefit from gpu parallelization and runs slightly faster on the cpu.</p>\n<p>The results from the FGS fit are then used to seed the fits to the AIRS data.  The model is fit to each of the 282 AIRS wavelengths separately, although these fits are all performed in parallel using torch's Adam minimizer.  The AIRS fits run for a fixed number of iterations, but terminate early if the chisquare function stops decreasing quickly enough.  In order to save time on the AIRS fits, the Tcenter and T parameters are fixed to the values obtained from the FGS fit; however, the tau, LDC, and transit depth parameters (along with all four background polynomial parameters) remain free and independent for each wavelength.</p>\n<p>Uncertainties for all parameters at each wavelength are estimated as the square root of the diagonal elements of the inverse of the Hessian matrix at the minimum.  Due to time constraints, it is not always possible to find a good enough minimum for this procedure to work; even in the best case, where every wavelength finds the correct minimum and a non-singular Hessian matrix, the computation is time-consuming enough that it creates tension with the runtime limits of the competition.  Thus, the uncertainty calculation is done only for every other wavelength, and missing values (whether from failed minimizations or deliberate skipping) are imputed (or extrapolated) from their nearest neighbors.  In the event that all fits for a given transit fail to produce uncertainty estimates, the analysis moves forward using the average uncertainty values observed in successful fits to the training data.</p>\n<p>The following two images show two of the fitted light curves for a typical transit, planet 189083129.   The first image shows the light curve for the FGS, and the second for one of the light curves from AIRS, wavelength 150.  Both are well-described by the fit.  The grey bands show the boundaries of the initial transit detection.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Fc04b787d3c4568c3c463159dc8c3a9d5%2Ffgs_light_curve_fit_planet_1890803129.png?generation=1758850162966272&amp;alt=media\" alt=\"Fit to FGS light curve for planet 189083129\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Ffb81891486462a609aaa66a2d654c533%2Fch0_fit_planet1890803129_visit0_wavelength150.png?generation=1758850257917836&amp;alt=media\" alt=\"Fit to light curve for AIRS wavelength 150 for planet 189083129\"></p>\n<p>The following two images show the corresponding fits for a pathological case with truncated out-of-transit regions, planet 158006264 visit 1:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2F78aff473df8d1c6ce8287291a1b338c6%2Ffgs_light_curve_fit_planet_158006264.png?generation=1758850499483929&amp;alt=media\" alt=\"Fit to FGS light curve for planet 158006264\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2F4f6a3b710ef1ca82c373a9240a0ca937%2Fch0_fit_planet158006264_visit1_wavelength150.png?generation=1758850555301923&amp;alt=media\" alt=\"Fit to light curve for AIRS wavelength 150 for planet 158006264\"></p>\n<p>There is no special handling in the fitter for these pathological cases, but in the few cases I've looked at, the description of the light curve is nevertheless reasonable if the fitter achieves good convergence. </p>\n<p>In the offline environment used for training, the fitter typically achieves an adequate but not fantastic fit quality, although there are some cases where the fit quality is worse.  The following figure shows the distribution of chisquare/ndof for all planets and all wavelengths in the training set:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Fd843d6cc55983f7a2f9e59cf186ffbb4%2Fchisq_ndof.png?generation=1758850810978620&amp;alt=media\" alt=\"\"></p>\n<p>I have not attempted to reproduce this plot in the Kaggle notebook environment.</p>\n<p>In the Kaggle notebook environment, using a P100 gpu, the fit to the FGS data (including its uncertainty analysis) typically takes about 2 or 3 seconds to process one transit's light curve, while the parallel fit to all AIRS light curves typically takes about 6 or 7 seconds.  These figures are averages; there is substantial variability among transits.</p>\n<h1>Neural nets for smoothing, limb darkening, and bias correction</h1>\n<p>Although the raw results from the fit include a prediction and uncertainty for each wavelength, they do not leverage the statistical power of the observations at neighboring wavelengths, nor do they account for effects like limb darkening.  Two neural nets are trained to handle these tasks.  Both have roughly the same architecture: they take as input a set of observations for each wavelength; an MLP maps these inputs to a length-64 feature vector for each wavelength, and then a two-layer, 1-dimensional convolutional network predicts a residual correction to the initial spectrum and its uncertainty estimates.  The two networks differ only in their inputs; one constitutes the \"main\" analysis and the other is a fallback in case of fit failure.  The main network takes as input several features for each wavelength:</p>\n<ul>\n<li>the transit depth estimate and its uncertainty, both from the Gaussian baseline</li>\n<li>the orbital parameters</li>\n<li>the mean of the light curve before the light curve is normalized to 1</li>\n<li>the best-fit values of the background polynomial parameters and their fit uncertainty estimates (based on the Hessian)</li>\n<li>residuals, i.e. <code>(fit-observation)/uncertainty</code> in the pre-transit and post-transit regions as well as in a few subsets of the in-transit region (regions re-defined based on the output of the fit)</li>\n<li>the average fitted background and signal in these regions (not including signal in the pre- and post-transit regions, since that is by definition zero)</li>\n<li>the total transit duration and its uncertainty (based on the FGS fit)</li>\n<li>several observables describing the location of the image of the star on the spatial directions of the FGS and AIRS detectors:  mean and width of the bright spot, averaged over time, and the corresponding standard deviations and ranges (maximum across time minus minimum across time), as well as the number of masked pixels in this transit</li>\n<li>the fitted values of the transit depth, tau, and LDC parameters, as well as their uncertainties at this wavelength</li>\n<li>the fitted values of the transit depth, tau, and LDC parameters and thier uncertainties, averaged over sliding windows of size 3, 7, and 47 on the frequency axis</li>\n</ul>\n<p>The main network predicts a residual correction to the fitted transit depth and its fit uncertainty.</p>\n<p>The fallback network mirrors the main network, but it takes a reduced set of inputs which eliminate any features derived from the fit, so it is left with only the depth estimates and uncertainties from the baseline analysis and the spatial observables; it predicts a residual correction to the Gaussian baseline estimates instead of the values from the fit.</p>\n<p>During training, the labels are smeared as a means of data augmentation.  Before processing each batch, a random offset is sampled for each wavelength from a Gaussian distribution centered at zero with width given by a multiple of the fit uncertainty.  The multiple is 5 at the beginning of training, and it attenuates by a factor of 0.9998 after each epoch.  Training runs for 150k epochs.  </p>\n<p>At the beginning of training, the predicted spectra often differ from the labels by several times the predicted uncertainty, leading to large gradients.  To stabilize training at the beginning of the run, the uncertainty from the fit (to which the residual correction predicted by the network is added) is artificially increased by a factor which starts at 1000 and decreases linearly to zero over the course of 1000 epochs. </p>\n<p>Both of these networks are trained using a five-fold cross-validation, and for each network, the residual corrections predicted by the five folds are ensembled using a simple arithmetic mean.</p>\n<p>The fallback network is deemed unreliable if the initial-guess widths of the ingress and egress ratio are very different from each other, or if the initial-guess ingress or egress region is too close to the beginning or end of the observation window.   It is deemed reliable otherwise.</p>\n<p>The result reported to the submission.csv for a given wavelength is the one from the main analysis unless the fallback is deemed reliable and any of the following conditions are met:</p>\n<ul>\n<li>the sum of the linear, quadratic, and cubic polynomial coefficients for that wavelength is larger than 0.07</li>\n<li>the chisquare/ndof for the fit at that wavelength is larger than 1.5</li>\n<li>the fitted value of the tau parameter is larger than 40</li>\n<li>the fit uncertainty on the tau parameter is larger than 5</li>\n<li>any of the residuals input to the main neural net has a magnitude larger than 1</li>\n<li>the detected start or end of transit is too close to the edge of the observation window.  (This condition is an oversight; it should have been removed in light of the very similar but not identical reliability cut for the fallback analysis.)</li>\n<li>the LDC parameter is less than 0.1</li>\n</ul>\n<p>These selections are applied on a wavelength-by-wavelength basis.  Most of these cuts have only a small effect on the training set; most are motivated by a desire that the main network should not be asked to make a prediction on regions of its input space where there was very little training data.</p>\n<p>The neural net operations described above typically take about 1.5-2 seconds per transit in timing tests in the Kaggle notebook environment with a P100 gpu.</p>\n<h1>Combination of multiple visits</h1>\n<p>In cases where a planetary system was observed in multiple transits, the combined prediction was an error-weighted average of the predictions from the two transits, taking half the difference between the central values of the two predictions as a systematic uncertainty added in quadrature to the uncertainty obtained from propagation of uncertainty.</p>\n<h1>An unexpected snag</h1>\n<p>Validation checks performed while porting the code to the Kaggle notebook environment turned up a curious difference in behavior for the fitting portion of this analysis:  the minimizers generally did not do as good a job finding minima in the Kaggle environment as they did in the offline environment.  In particular, the chisquare/ndof values from the fits in the Kaggle environment were still generally ok, but they weren't quite as good as the were offline, and the corresponding uncertainty estimates were often a bit bigger.  As this issue only presented itself very close to the close of the competition, I did not manage to pin down the exact reason why.  I wasn't able to spot a difference in the inputs or the libraries that I know were used (I checked pytorch, numpy, scipy, and cuda/cudnn versions, but not other supporting libraries.)  I was able to mitigate this effect a bit by adding checks to the notebook code that would trigger another call to the fitter in order to improve convergence when needed, but I wasn't able to convince myself that the problem was really solved.  I suspect this may be part of why there is a large gap between my local cross-validation scores and the actual leaderboard score that I got in the end.  If anyone ran into similar difficulties, I'd be curious to hear about it.</p>\n<h1>Code availability</h1>\n<p>I have attached to this writeup a link to the submission notebook.  Training code and model weights are available upon request, but I think that the parts anyone might be interested in and able to reuse elsewhere will likely be in the notebook.</p>",
  "messages": [
    {
      "id": 3294359,
      "postDate": "2025-09-26T01:04:11.580Z",
      "content": "<h1>Executive summary</h1>\n<p>A chisquare fit of a parametric function (3- or 5-parameter trapezoid for signal, cubic polynomial for background) to all wavelengths provides inputs to a small, 1D convolutional neural net which smooths the fit results and attempts to correct for limb darkening and any internal fit bias which may be present. The most significant challenge to this approach was getting the fits to achieve adequate convergence within the runtime constraints of the competition, and so this write-up emphasizes timing measurements and time-saving techniques where appropriate.</p>\n<p>A potentially interesting byproduct of this analysis is a wavelength-by-wavelength measurement of ingress/egress time and limb darkening; however, as no ground truth was provided for these quantities, it is difficult to be sure that these measurements, as they are now, are fit for any downstream purpose.</p>\n<h1>Preprocessing</h1>\n<p>The initial processing and calibration of the raw data follows the standard procedure, described in <a href=\"https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data\" target=\"_blank\">this notebook</a>; however, the implementation used in the present solution performs several of these operations on the gpu in order to save time.</p>\n<p>A simple outlier-exclusion procedure is applied to both the FGS and airs data.  Before summing over spatial dimensions, the mean and standard deviation of the readout for each pixel are computed in a sliding window over 50 timesteps. Any time a pixel's readout is more than 3 sigma above or below the mean in that window, its value is replaced by the mean.</p>\n<p>Both the AIRS and FGS data are summed over their respective spatial dimensions to obtain one light curve for each wavelength. FGS light curve are summed in bins of 12 timesteps to match the AIRS light curves, and then both are collected into bins of 15 timesteps, so that each light curve for each detector and wavelength is comprised of 375 time bins.  Each light curve is normalized so that its average over time is 1.  The mean and standard error of the observations in each time bin provide an initial set of y and y uncertainty values for the fit.  In an attempt to improve the convergence of the fit, the y uncertainty values are smoothed by replacing each y uncertainty value with the mean of the uncertainties in a sliding window of 5 time bins.  There was unfortunately not enough time to perform an ablation study to verify the impact of this step on the final version of the analysis before the close of the competition. The x values used in the fit are evenly-spaced numbers ranging from 0 at the start of data collection to 1 at the end.</p>\n<p>An initial detection of the ingress and egress regions is obtained by performing a continuous wavelet transform (using the 'cgau1' wavelet in pywt) of the FGS light curve and searching for maxima in the resulting wavelet coefficients.  The function that does this returns a mean and a width for the two largest peaks in the transformed data.  The initial estimate of the ingress region is taken to be a window of +/- 2 sigma around the earlier of the two peaks, and the intial estimate for the egress region is taken to be a window of +/- 2sigma around the other peak.  The in-transit region is the region between the ingress and egress regions, and the out-of-transit region is the rest of the light curve.  These initial estimates serve two purposes in this analysis:  they are used to compute initial guesses for certain fit parameters, and they are also used in a \"Gaussian baseline\" analysis that is used as a fallback when there are indications that a fit has failed.</p>\n<p>The \"Gaussian baseline\" is a simple cut-and-count estimate of the transit depth.  Given a normalized light curve (from either FGS or from any AIRS wavelength) and the initial estimates of the ingress and egress regions, the baseline analysis computes the fit depth as the difference between the mean light curve value in the in-transit and out-of-transit regions, divided by the mean light curve value in the out-of-transit region.  Gaussian propagation of uncertainty provides a crude estimate of the uncertainty on that fit depth.  This estimate is used later to compute features that are included in the neural net inputs, and it is also used as a replacement for the fit values themselves in a \"fallback\" result that is reported in the submission when the fit results are deemed too unreliable (according to a set of selections defined below).</p>\n<p>The initial estimates of the in-transit and out-of-transit regions are also used to compute initial values for the parameters of the FGS fit described below. The distance between the means is taken as an initial guess for T in the fit; the midpoint is taken as an initial guess for Tcenter; and the average of the two widths is the initial guess for tau.  An initial guess for the limb-darkening parameter is derived from the range of normalized light-curve intensities present in the in-transit region.  The background polynomial parameters are initialized such that the constant term is 1 and the linear, quadratic, and cubic terms are all 0.  The initial guess for the fit depth is taken from the Gaussian baseline.  Initial guesses for the AIRS fits are taken from the output of the FGS fit.</p>\n<p>Timing estimates of the preprocessing phase of the analysis include all of the above operations and are performed in the Kaggle notebook environment with a P100 gpu, usually running on 50-100 examples from the training set.  There was some variability between one timing run and the next, with this phase of the analysis typically clocking in at anywhere between 12 and 15 seconds per transit on average. </p>\n<h1>Parametric fit</h1>\n<p>The signal is modeled as a trapezoid, multiplied by a parabola centered at the trapezoid's midpoint.  Its shape is characterized by five parameters:</p>\n<ul>\n<li>Tcenter, the midpoint of the trapezoid and parabola.</li>\n<li>T, the duration of the transit, i.e. the time between the middle of the ingress (left trapezoid leg) and the middle of the egress (right trapezoid leg)</li>\n<li>tau, the ingress/egress time.  This parameter is shared between the ingress and egress regions; that is, the signal model does not allow for the ingress and egress times to be different.</li>\n<li>LDC, the strength of the parabola which models the effect of limb darkening.  Note that this is not the same as any of the parameters of limb darkening models in the literature; it is more properly understood as an empirical parameter which describes the shape of the top of the trapezoid.</li>\n</ul>\n<p>The values of these parameters in the fit are all limited to fixed ranges by transforming the fit parameters through tanh or sigmoid functions. The trapezoid shape is implemented in a fast, gpu-friendly way using pytorch  (code simplified slightly for readability):</p>\n<p><code>trapezoid=torch.clip((torch.arange(0,n_timesteps)-tstart)/tau,0,1)*torch.clip((tend-torch.arange(0,n_timesteps))/tau,0,1)</code></p>\n<p>where <code>tstart</code> and <code>tend</code> are <code>Tcenter-T/2</code> and <code>Tcenter+T/2</code>, respectively.  The limb-darkening effect is applied as a multiplicative correction to this:</p>\n<pre><code>=torch.es((n_timesteps,))\n=(torch.arange(,n_timesteps)-Tcenter)/max(T/,)\n=torch.clip(limbdark-ldc1*limbdark_x**,,)\n=limbdark*trapezoid\n</code></pre>\n<p>This four-parameter shape is normalized so that its maximum value is 1 and is then combined with the fourth-order polynomial background with the help of another fit parameter, the transit depth:</p>\n<p><code>fit_y=fit_y*(1-fit_depth*signal_shape)</code></p>\n<p>The transit depth is constrained to be positive by transforming the corresponding fit parameter using a softplus function.  Since the model is implemented in pytorch, it is straightforward to provide a gradient estimate to the minimizer using torch's automatic differentiation.  A chisquare function is defined in the usual way as the sum of <code>(obs_y-fit_y)**2/y_err**2</code> over all time and wavelength bins.  The model is first fit to the FGS data using scipy.optimize.minimize() to find the fit parameters that minimize chisquare, allowing all parameters to float freely in the fit.  Up to three calls to minimize() are allowed:  the first uses the default minimizer, BFGS; if that fails (or if the corresponding uncertainty analysis, described below, fails), then a second call is made to the SLSQP minimizer; if that fails, then a third call is made to the Powell minimzer.  Since the FGS data consist of only one wavelength, i.e. only one light curve, this fit does not benefit from gpu parallelization and runs slightly faster on the cpu.</p>\n<p>The results from the FGS fit are then used to seed the fits to the AIRS data.  The model is fit to each of the 282 AIRS wavelengths separately, although these fits are all performed in parallel using torch's Adam minimizer.  The AIRS fits run for a fixed number of iterations, but terminate early if the chisquare function stops decreasing quickly enough.  In order to save time on the AIRS fits, the Tcenter and T parameters are fixed to the values obtained from the FGS fit; however, the tau, LDC, and transit depth parameters (along with all four background polynomial parameters) remain free and independent for each wavelength.</p>\n<p>Uncertainties for all parameters at each wavelength are estimated as the square root of the diagonal elements of the inverse of the Hessian matrix at the minimum.  Due to time constraints, it is not always possible to find a good enough minimum for this procedure to work; even in the best case, where every wavelength finds the correct minimum and a non-singular Hessian matrix, the computation is time-consuming enough that it creates tension with the runtime limits of the competition.  Thus, the uncertainty calculation is done only for every other wavelength, and missing values (whether from failed minimizations or deliberate skipping) are imputed (or extrapolated) from their nearest neighbors.  In the event that all fits for a given transit fail to produce uncertainty estimates, the analysis moves forward using the average uncertainty values observed in successful fits to the training data.</p>\n<p>The following two images show two of the fitted light curves for a typical transit, planet 189083129.   The first image shows the light curve for the FGS, and the second for one of the light curves from AIRS, wavelength 150.  Both are well-described by the fit.  The grey bands show the boundaries of the initial transit detection.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Fc04b787d3c4568c3c463159dc8c3a9d5%2Ffgs_light_curve_fit_planet_1890803129.png?generation=1758850162966272&amp;alt=media\" alt=\"Fit to FGS light curve for planet 189083129\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Ffb81891486462a609aaa66a2d654c533%2Fch0_fit_planet1890803129_visit0_wavelength150.png?generation=1758850257917836&amp;alt=media\" alt=\"Fit to light curve for AIRS wavelength 150 for planet 189083129\"></p>\n<p>The following two images show the corresponding fits for a pathological case with truncated out-of-transit regions, planet 158006264 visit 1:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2F78aff473df8d1c6ce8287291a1b338c6%2Ffgs_light_curve_fit_planet_158006264.png?generation=1758850499483929&amp;alt=media\" alt=\"Fit to FGS light curve for planet 158006264\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2F4f6a3b710ef1ca82c373a9240a0ca937%2Fch0_fit_planet158006264_visit1_wavelength150.png?generation=1758850555301923&amp;alt=media\" alt=\"Fit to light curve for AIRS wavelength 150 for planet 158006264\"></p>\n<p>There is no special handling in the fitter for these pathological cases, but in the few cases I've looked at, the description of the light curve is nevertheless reasonable if the fitter achieves good convergence. </p>\n<p>In the offline environment used for training, the fitter typically achieves an adequate but not fantastic fit quality, although there are some cases where the fit quality is worse.  The following figure shows the distribution of chisquare/ndof for all planets and all wavelengths in the training set:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Fd843d6cc55983f7a2f9e59cf186ffbb4%2Fchisq_ndof.png?generation=1758850810978620&amp;alt=media\" alt=\"\"></p>\n<p>I have not attempted to reproduce this plot in the Kaggle notebook environment.</p>\n<p>In the Kaggle notebook environment, using a P100 gpu, the fit to the FGS data (including its uncertainty analysis) typically takes about 2 or 3 seconds to process one transit's light curve, while the parallel fit to all AIRS light curves typically takes about 6 or 7 seconds.  These figures are averages; there is substantial variability among transits.</p>\n<h1>Neural nets for smoothing, limb darkening, and bias correction</h1>\n<p>Although the raw results from the fit include a prediction and uncertainty for each wavelength, they do not leverage the statistical power of the observations at neighboring wavelengths, nor do they account for effects like limb darkening.  Two neural nets are trained to handle these tasks.  Both have roughly the same architecture: they take as input a set of observations for each wavelength; an MLP maps these inputs to a length-64 feature vector for each wavelength, and then a two-layer, 1-dimensional convolutional network predicts a residual correction to the initial spectrum and its uncertainty estimates.  The two networks differ only in their inputs; one constitutes the \"main\" analysis and the other is a fallback in case of fit failure.  The main network takes as input several features for each wavelength:</p>\n<ul>\n<li>the transit depth estimate and its uncertainty, both from the Gaussian baseline</li>\n<li>the orbital parameters</li>\n<li>the mean of the light curve before the light curve is normalized to 1</li>\n<li>the best-fit values of the background polynomial parameters and their fit uncertainty estimates (based on the Hessian)</li>\n<li>residuals, i.e. <code>(fit-observation)/uncertainty</code> in the pre-transit and post-transit regions as well as in a few subsets of the in-transit region (regions re-defined based on the output of the fit)</li>\n<li>the average fitted background and signal in these regions (not including signal in the pre- and post-transit regions, since that is by definition zero)</li>\n<li>the total transit duration and its uncertainty (based on the FGS fit)</li>\n<li>several observables describing the location of the image of the star on the spatial directions of the FGS and AIRS detectors:  mean and width of the bright spot, averaged over time, and the corresponding standard deviations and ranges (maximum across time minus minimum across time), as well as the number of masked pixels in this transit</li>\n<li>the fitted values of the transit depth, tau, and LDC parameters, as well as their uncertainties at this wavelength</li>\n<li>the fitted values of the transit depth, tau, and LDC parameters and thier uncertainties, averaged over sliding windows of size 3, 7, and 47 on the frequency axis</li>\n</ul>\n<p>The main network predicts a residual correction to the fitted transit depth and its fit uncertainty.</p>\n<p>The fallback network mirrors the main network, but it takes a reduced set of inputs which eliminate any features derived from the fit, so it is left with only the depth estimates and uncertainties from the baseline analysis and the spatial observables; it predicts a residual correction to the Gaussian baseline estimates instead of the values from the fit.</p>\n<p>During training, the labels are smeared as a means of data augmentation.  Before processing each batch, a random offset is sampled for each wavelength from a Gaussian distribution centered at zero with width given by a multiple of the fit uncertainty.  The multiple is 5 at the beginning of training, and it attenuates by a factor of 0.9998 after each epoch.  Training runs for 150k epochs.  </p>\n<p>At the beginning of training, the predicted spectra often differ from the labels by several times the predicted uncertainty, leading to large gradients.  To stabilize training at the beginning of the run, the uncertainty from the fit (to which the residual correction predicted by the network is added) is artificially increased by a factor which starts at 1000 and decreases linearly to zero over the course of 1000 epochs. </p>\n<p>Both of these networks are trained using a five-fold cross-validation, and for each network, the residual corrections predicted by the five folds are ensembled using a simple arithmetic mean.</p>\n<p>The fallback network is deemed unreliable if the initial-guess widths of the ingress and egress ratio are very different from each other, or if the initial-guess ingress or egress region is too close to the beginning or end of the observation window.   It is deemed reliable otherwise.</p>\n<p>The result reported to the submission.csv for a given wavelength is the one from the main analysis unless the fallback is deemed reliable and any of the following conditions are met:</p>\n<ul>\n<li>the sum of the linear, quadratic, and cubic polynomial coefficients for that wavelength is larger than 0.07</li>\n<li>the chisquare/ndof for the fit at that wavelength is larger than 1.5</li>\n<li>the fitted value of the tau parameter is larger than 40</li>\n<li>the fit uncertainty on the tau parameter is larger than 5</li>\n<li>any of the residuals input to the main neural net has a magnitude larger than 1</li>\n<li>the detected start or end of transit is too close to the edge of the observation window.  (This condition is an oversight; it should have been removed in light of the very similar but not identical reliability cut for the fallback analysis.)</li>\n<li>the LDC parameter is less than 0.1</li>\n</ul>\n<p>These selections are applied on a wavelength-by-wavelength basis.  Most of these cuts have only a small effect on the training set; most are motivated by a desire that the main network should not be asked to make a prediction on regions of its input space where there was very little training data.</p>\n<p>The neural net operations described above typically take about 1.5-2 seconds per transit in timing tests in the Kaggle notebook environment with a P100 gpu.</p>\n<h1>Combination of multiple visits</h1>\n<p>In cases where a planetary system was observed in multiple transits, the combined prediction was an error-weighted average of the predictions from the two transits, taking half the difference between the central values of the two predictions as a systematic uncertainty added in quadrature to the uncertainty obtained from propagation of uncertainty.</p>\n<h1>An unexpected snag</h1>\n<p>Validation checks performed while porting the code to the Kaggle notebook environment turned up a curious difference in behavior for the fitting portion of this analysis:  the minimizers generally did not do as good a job finding minima in the Kaggle environment as they did in the offline environment.  In particular, the chisquare/ndof values from the fits in the Kaggle environment were still generally ok, but they weren't quite as good as the were offline, and the corresponding uncertainty estimates were often a bit bigger.  As this issue only presented itself very close to the close of the competition, I did not manage to pin down the exact reason why.  I wasn't able to spot a difference in the inputs or the libraries that I know were used (I checked pytorch, numpy, scipy, and cuda/cudnn versions, but not other supporting libraries.)  I was able to mitigate this effect a bit by adding checks to the notebook code that would trigger another call to the fitter in order to improve convergence when needed, but I wasn't able to convince myself that the problem was really solved.  I suspect this may be part of why there is a large gap between my local cross-validation scores and the actual leaderboard score that I got in the end.  If anyone ran into similar difficulties, I'd be curious to hear about it.</p>\n<h1>Code availability</h1>\n<p>I have attached to this writeup a link to the submission notebook.  Training code and model weights are available upon request, but I think that the parts anyone might be interested in and able to reuse elsewhere will likely be in the notebook.</p>",
      "rawMarkdown": "# Executive summary\n\nA chisquare fit of a parametric function (3- or 5-parameter trapezoid for signal, cubic polynomial for background) to all wavelengths provides inputs to a small, 1D convolutional neural net which smooths the fit results and attempts to correct for limb darkening and any internal fit bias which may be present. The most significant challenge to this approach was getting the fits to achieve adequate convergence within the runtime constraints of the competition, and so this write-up emphasizes timing measurements and time-saving techniques where appropriate.\n\nA potentially interesting byproduct of this analysis is a wavelength-by-wavelength measurement of ingress/egress time and limb darkening; however, as no ground truth was provided for these quantities, it is difficult to be sure that these measurements, as they are now, are fit for any downstream purpose.\n  \n# Preprocessing\n  \nThe initial processing and calibration of the raw data follows the standard procedure, described in [this notebook](https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data); however, the implementation used in the present solution performs several of these operations on the gpu in order to save time.\n  \nA simple outlier-exclusion procedure is applied to both the FGS and airs data.  Before summing over spatial dimensions, the mean and standard deviation of the readout for each pixel are computed in a sliding window over 50 timesteps. Any time a pixel's readout is more than 3 sigma above or below the mean in that window, its value is replaced by the mean.\n\nBoth the AIRS and FGS data are summed over their respective spatial dimensions to obtain one light curve for each wavelength. FGS light curve are summed in bins of 12 timesteps to match the AIRS light curves, and then both are collected into bins of 15 timesteps, so that each light curve for each detector and wavelength is comprised of 375 time bins.  Each light curve is normalized so that its average over time is 1.  The mean and standard error of the observations in each time bin provide an initial set of y and y uncertainty values for the fit.  In an attempt to improve the convergence of the fit, the y uncertainty values are smoothed by replacing each y uncertainty value with the mean of the uncertainties in a sliding window of 5 time bins.  There was unfortunately not enough time to perform an ablation study to verify the impact of this step on the final version of the analysis before the close of the competition. The x values used in the fit are evenly-spaced numbers ranging from 0 at the start of data collection to 1 at the end.\n\nAn initial detection of the ingress and egress regions is obtained by performing a continuous wavelet transform (using the 'cgau1' wavelet in pywt) of the FGS light curve and searching for maxima in the resulting wavelet coefficients.  The function that does this returns a mean and a width for the two largest peaks in the transformed data.  The initial estimate of the ingress region is taken to be a window of +/- 2 sigma around the earlier of the two peaks, and the intial estimate for the egress region is taken to be a window of +/- 2sigma around the other peak.  The in-transit region is the region between the ingress and egress regions, and the out-of-transit region is the rest of the light curve.  These initial estimates serve two purposes in this analysis:  they are used to compute initial guesses for certain fit parameters, and they are also used in a \"Gaussian baseline\" analysis that is used as a fallback when there are indications that a fit has failed.\n    \nThe \"Gaussian baseline\" is a simple cut-and-count estimate of the transit depth.  Given a normalized light curve (from either FGS or from any AIRS wavelength) and the initial estimates of the ingress and egress regions, the baseline analysis computes the fit depth as the difference between the mean light curve value in the in-transit and out-of-transit regions, divided by the mean light curve value in the out-of-transit region.  Gaussian propagation of uncertainty provides a crude estimate of the uncertainty on that fit depth.  This estimate is used later to compute features that are included in the neural net inputs, and it is also used as a replacement for the fit values themselves in a \"fallback\" result that is reported in the submission when the fit results are deemed too unreliable (according to a set of selections defined below).\n\nThe initial estimates of the in-transit and out-of-transit regions are also used to compute initial values for the parameters of the FGS fit described below. The distance between the means is taken as an initial guess for T in the fit; the midpoint is taken as an initial guess for Tcenter; and the average of the two widths is the initial guess for tau.  An initial guess for the limb-darkening parameter is derived from the range of normalized light-curve intensities present in the in-transit region.  The background polynomial parameters are initialized such that the constant term is 1 and the linear, quadratic, and cubic terms are all 0.  The initial guess for the fit depth is taken from the Gaussian baseline.  Initial guesses for the AIRS fits are taken from the output of the FGS fit.\n\nTiming estimates of the preprocessing phase of the analysis include all of the above operations and are performed in the Kaggle notebook environment with a P100 gpu, usually running on 50-100 examples from the training set.  There was some variability between one timing run and the next, with this phase of the analysis typically clocking in at anywhere between 12 and 15 seconds per transit on average. \n\n# Parametric fit\n\nThe signal is modeled as a trapezoid, multiplied by a parabola centered at the trapezoid's midpoint.  Its shape is characterized by five parameters:\n * Tcenter, the midpoint of the trapezoid and parabola.\n * T, the duration of the transit, i.e. the time between the middle of the ingress (left trapezoid leg) and the middle of the egress (right trapezoid leg)\n * tau, the ingress/egress time.  This parameter is shared between the ingress and egress regions; that is, the signal model does not allow for the ingress and egress times to be different.\n * LDC, the strength of the parabola which models the effect of limb darkening.  Note that this is not the same as any of the parameters of limb darkening models in the literature; it is more properly understood as an empirical parameter which describes the shape of the top of the trapezoid.\n\nThe values of these parameters in the fit are all limited to fixed ranges by transforming the fit parameters through tanh or sigmoid functions. The trapezoid shape is implemented in a fast, gpu-friendly way using pytorch  (code simplified slightly for readability):\n\n  `trapezoid=torch.clip((torch.arange(0,n_timesteps)-tstart)/tau,0,1)*torch.clip((tend-torch.arange(0,n_timesteps))/tau,0,1)`\n\nwhere `tstart` and `tend` are `Tcenter-T/2` and `Tcenter+T/2`, respectively.  The limb-darkening effect is applied as a multiplicative correction to this:\n\n```\nlimbdark=torch.ones((n_timesteps,))\nlimbdark_x=(torch.arange(0,n_timesteps)-Tcenter)/max(T/2,1)\nlimbdark=torch.clip(limbdark-ldc1*limbdark_x**2,0,1)\nsignal_shape=limbdark*trapezoid\n```\n\nThis four-parameter shape is normalized so that its maximum value is 1 and is then combined with the fourth-order polynomial background with the help of another fit parameter, the transit depth:\n\n`fit_y=fit_y*(1-fit_depth*signal_shape)`\n\nThe transit depth is constrained to be positive by transforming the corresponding fit parameter using a softplus function.  Since the model is implemented in pytorch, it is straightforward to provide a gradient estimate to the minimizer using torch's automatic differentiation.  A chisquare function is defined in the usual way as the sum of `(obs_y-fit_y)**2/y_err**2` over all time and wavelength bins.  The model is first fit to the FGS data using scipy.optimize.minimize() to find the fit parameters that minimize chisquare, allowing all parameters to float freely in the fit.  Up to three calls to minimize() are allowed:  the first uses the default minimizer, BFGS; if that fails (or if the corresponding uncertainty analysis, described below, fails), then a second call is made to the SLSQP minimizer; if that fails, then a third call is made to the Powell minimzer.  Since the FGS data consist of only one wavelength, i.e. only one light curve, this fit does not benefit from gpu parallelization and runs slightly faster on the cpu.\n\nThe results from the FGS fit are then used to seed the fits to the AIRS data.  The model is fit to each of the 282 AIRS wavelengths separately, although these fits are all performed in parallel using torch's Adam minimizer.  The AIRS fits run for a fixed number of iterations, but terminate early if the chisquare function stops decreasing quickly enough.  In order to save time on the AIRS fits, the Tcenter and T parameters are fixed to the values obtained from the FGS fit; however, the tau, LDC, and transit depth parameters (along with all four background polynomial parameters) remain free and independent for each wavelength.\n\nUncertainties for all parameters at each wavelength are estimated as the square root of the diagonal elements of the inverse of the Hessian matrix at the minimum.  Due to time constraints, it is not always possible to find a good enough minimum for this procedure to work; even in the best case, where every wavelength finds the correct minimum and a non-singular Hessian matrix, the computation is time-consuming enough that it creates tension with the runtime limits of the competition.  Thus, the uncertainty calculation is done only for every other wavelength, and missing values (whether from failed minimizations or deliberate skipping) are imputed (or extrapolated) from their nearest neighbors.  In the event that all fits for a given transit fail to produce uncertainty estimates, the analysis moves forward using the average uncertainty values observed in successful fits to the training data.\n\nThe following two images show two of the fitted light curves for a typical transit, planet 189083129.   The first image shows the light curve for the FGS, and the second for one of the light curves from AIRS, wavelength 150.  Both are well-described by the fit.  The grey bands show the boundaries of the initial transit detection.\n\n![Fit to FGS light curve for planet 189083129](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Fc04b787d3c4568c3c463159dc8c3a9d5%2Ffgs_light_curve_fit_planet_1890803129.png?generation=1758850162966272&alt=media)\n\n![Fit to light curve for AIRS wavelength 150 for planet 189083129](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Ffb81891486462a609aaa66a2d654c533%2Fch0_fit_planet1890803129_visit0_wavelength150.png?generation=1758850257917836&alt=media)\n\nThe following two images show the corresponding fits for a pathological case with truncated out-of-transit regions, planet 158006264 visit 1:\n\n![Fit to FGS light curve for planet 158006264](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2F78aff473df8d1c6ce8287291a1b338c6%2Ffgs_light_curve_fit_planet_158006264.png?generation=1758850499483929&alt=media)\n\n![Fit to light curve for AIRS wavelength 150 for planet 158006264](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2F4f6a3b710ef1ca82c373a9240a0ca937%2Fch0_fit_planet158006264_visit1_wavelength150.png?generation=1758850555301923&alt=media)\n\nThere is no special handling in the fitter for these pathological cases, but in the few cases I've looked at, the description of the light curve is nevertheless reasonable if the fitter achieves good convergence. \n\nIn the offline environment used for training, the fitter typically achieves an adequate but not fantastic fit quality, although there are some cases where the fit quality is worse.  The following figure shows the distribution of chisquare/ndof for all planets and all wavelengths in the training set:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Fd843d6cc55983f7a2f9e59cf186ffbb4%2Fchisq_ndof.png?generation=1758850810978620&alt=media)\n\nI have not attempted to reproduce this plot in the Kaggle notebook environment.\n\nIn the Kaggle notebook environment, using a P100 gpu, the fit to the FGS data (including its uncertainty analysis) typically takes about 2 or 3 seconds to process one transit's light curve, while the parallel fit to all AIRS light curves typically takes about 6 or 7 seconds.  These figures are averages; there is substantial variability among transits.\n\n#Neural nets for smoothing, limb darkening, and bias correction\n\nAlthough the raw results from the fit include a prediction and uncertainty for each wavelength, they do not leverage the statistical power of the observations at neighboring wavelengths, nor do they account for effects like limb darkening.  Two neural nets are trained to handle these tasks.  Both have roughly the same architecture: they take as input a set of observations for each wavelength; an MLP maps these inputs to a length-64 feature vector for each wavelength, and then a two-layer, 1-dimensional convolutional network predicts a residual correction to the initial spectrum and its uncertainty estimates.  The two networks differ only in their inputs; one constitutes the \"main\" analysis and the other is a fallback in case of fit failure.  The main network takes as input several features for each wavelength:\n  * the transit depth estimate and its uncertainty, both from the Gaussian baseline\n  * the orbital parameters\n  * the mean of the light curve before the light curve is normalized to 1\n  * the best-fit values of the background polynomial parameters and their fit uncertainty estimates (based on the Hessian)\n  * residuals, i.e. `(fit-observation)/uncertainty` in the pre-transit and post-transit regions as well as in a few subsets of the in-transit region (regions re-defined based on the output of the fit)\n  * the average fitted background and signal in these regions (not including signal in the pre- and post-transit regions, since that is by definition zero)\n  * the total transit duration and its uncertainty (based on the FGS fit)\n  * several observables describing the location of the image of the star on the spatial directions of the FGS and AIRS detectors:  mean and width of the bright spot, averaged over time, and the corresponding standard deviations and ranges (maximum across time minus minimum across time), as well as the number of masked pixels in this transit\n  * the fitted values of the transit depth, tau, and LDC parameters, as well as their uncertainties at this wavelength\n  * the fitted values of the transit depth, tau, and LDC parameters and thier uncertainties, averaged over sliding windows of size 3, 7, and 47 on the frequency axis\n\nThe main network predicts a residual correction to the fitted transit depth and its fit uncertainty.\n\nThe fallback network mirrors the main network, but it takes a reduced set of inputs which eliminate any features derived from the fit, so it is left with only the depth estimates and uncertainties from the baseline analysis and the spatial observables; it predicts a residual correction to the Gaussian baseline estimates instead of the values from the fit.\n\nDuring training, the labels are smeared as a means of data augmentation.  Before processing each batch, a random offset is sampled for each wavelength from a Gaussian distribution centered at zero with width given by a multiple of the fit uncertainty.  The multiple is 5 at the beginning of training, and it attenuates by a factor of 0.9998 after each epoch.  Training runs for 150k epochs.  \n\nAt the beginning of training, the predicted spectra often differ from the labels by several times the predicted uncertainty, leading to large gradients.  To stabilize training at the beginning of the run, the uncertainty from the fit (to which the residual correction predicted by the network is added) is artificially increased by a factor which starts at 1000 and decreases linearly to zero over the course of 1000 epochs. \n\nBoth of these networks are trained using a five-fold cross-validation, and for each network, the residual corrections predicted by the five folds are ensembled using a simple arithmetic mean.\n\nThe fallback network is deemed unreliable if the initial-guess widths of the ingress and egress ratio are very different from each other, or if the initial-guess ingress or egress region is too close to the beginning or end of the observation window.   It is deemed reliable otherwise.\n\nThe result reported to the submission.csv for a given wavelength is the one from the main analysis unless the fallback is deemed reliable and any of the following conditions are met:\n  * the sum of the linear, quadratic, and cubic polynomial coefficients for that wavelength is larger than 0.07\n  * the chisquare/ndof for the fit at that wavelength is larger than 1.5\n  * the fitted value of the tau parameter is larger than 40\n  * the fit uncertainty on the tau parameter is larger than 5\n  * any of the residuals input to the main neural net has a magnitude larger than 1\n  * the detected start or end of transit is too close to the edge of the observation window.  (This condition is an oversight; it should have been removed in light of the very similar but not identical reliability cut for the fallback analysis.)\n  * the LDC parameter is less than 0.1\n\nThese selections are applied on a wavelength-by-wavelength basis.  Most of these cuts have only a small effect on the training set; most are motivated by a desire that the main network should not be asked to make a prediction on regions of its input space where there was very little training data.\n\nThe neural net operations described above typically take about 1.5-2 seconds per transit in timing tests in the Kaggle notebook environment with a P100 gpu.\n\n# Combination of multiple visits\n\nIn cases where a planetary system was observed in multiple transits, the combined prediction was an error-weighted average of the predictions from the two transits, taking half the difference between the central values of the two predictions as a systematic uncertainty added in quadrature to the uncertainty obtained from propagation of uncertainty.\n\n# An unexpected snag\n\nValidation checks performed while porting the code to the Kaggle notebook environment turned up a curious difference in behavior for the fitting portion of this analysis:  the minimizers generally did not do as good a job finding minima in the Kaggle environment as they did in the offline environment.  In particular, the chisquare/ndof values from the fits in the Kaggle environment were still generally ok, but they weren't quite as good as the were offline, and the corresponding uncertainty estimates were often a bit bigger.  As this issue only presented itself very close to the close of the competition, I did not manage to pin down the exact reason why.  I wasn't able to spot a difference in the inputs or the libraries that I know were used (I checked pytorch, numpy, scipy, and cuda/cudnn versions, but not other supporting libraries.)  I was able to mitigate this effect a bit by adding checks to the notebook code that would trigger another call to the fitter in order to improve convergence when needed, but I wasn't able to convince myself that the problem was really solved.  I suspect this may be part of why there is a large gap between my local cross-validation scores and the actual leaderboard score that I got in the end.  If anyone ran into similar difficulties, I'd be curious to hear about it.\n\n# Code availability\n\nI have attached to this writeup a link to the submission notebook.  Training code and model weights are available upon request, but I think that the parts anyone might be interested in and able to reuse elsewhere will likely be in the notebook.\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3294359": "# Executive summary\n\nA chisquare fit of a parametric function (3- or 5-parameter trapezoid for signal, cubic polynomial for background) to all wavelengths provides inputs to a small, 1D convolutional neural net which smooths the fit results and attempts to correct for limb darkening and any internal fit bias which may be present. The most significant challenge to this approach was getting the fits to achieve adequate convergence within the runtime constraints of the competition, and so this write-up emphasizes timing measurements and time-saving techniques where appropriate.\n\nA potentially interesting byproduct of this analysis is a wavelength-by-wavelength measurement of ingress/egress time and limb darkening; however, as no ground truth was provided for these quantities, it is difficult to be sure that these measurements, as they are now, are fit for any downstream purpose.\n  \n# Preprocessing\n  \nThe initial processing and calibration of the raw data follows the standard procedure, described in [this notebook](https://www.kaggle.com/code/gordonyip/calibrating-and-binning-ariel-data); however, the implementation used in the present solution performs several of these operations on the gpu in order to save time.\n  \nA simple outlier-exclusion procedure is applied to both the FGS and airs data.  Before summing over spatial dimensions, the mean and standard deviation of the readout for each pixel are computed in a sliding window over 50 timesteps. Any time a pixel's readout is more than 3 sigma above or below the mean in that window, its value is replaced by the mean.\n\nBoth the AIRS and FGS data are summed over their respective spatial dimensions to obtain one light curve for each wavelength. FGS light curve are summed in bins of 12 timesteps to match the AIRS light curves, and then both are collected into bins of 15 timesteps, so that each light curve for each detector and wavelength is comprised of 375 time bins.  Each light curve is normalized so that its average over time is 1.  The mean and standard error of the observations in each time bin provide an initial set of y and y uncertainty values for the fit.  In an attempt to improve the convergence of the fit, the y uncertainty values are smoothed by replacing each y uncertainty value with the mean of the uncertainties in a sliding window of 5 time bins.  There was unfortunately not enough time to perform an ablation study to verify the impact of this step on the final version of the analysis before the close of the competition. The x values used in the fit are evenly-spaced numbers ranging from 0 at the start of data collection to 1 at the end.\n\nAn initial detection of the ingress and egress regions is obtained by performing a continuous wavelet transform (using the 'cgau1' wavelet in pywt) of the FGS light curve and searching for maxima in the resulting wavelet coefficients.  The function that does this returns a mean and a width for the two largest peaks in the transformed data.  The initial estimate of the ingress region is taken to be a window of +/- 2 sigma around the earlier of the two peaks, and the intial estimate for the egress region is taken to be a window of +/- 2sigma around the other peak.  The in-transit region is the region between the ingress and egress regions, and the out-of-transit region is the rest of the light curve.  These initial estimates serve two purposes in this analysis:  they are used to compute initial guesses for certain fit parameters, and they are also used in a \"Gaussian baseline\" analysis that is used as a fallback when there are indications that a fit has failed.\n    \nThe \"Gaussian baseline\" is a simple cut-and-count estimate of the transit depth.  Given a normalized light curve (from either FGS or from any AIRS wavelength) and the initial estimates of the ingress and egress regions, the baseline analysis computes the fit depth as the difference between the mean light curve value in the in-transit and out-of-transit regions, divided by the mean light curve value in the out-of-transit region.  Gaussian propagation of uncertainty provides a crude estimate of the uncertainty on that fit depth.  This estimate is used later to compute features that are included in the neural net inputs, and it is also used as a replacement for the fit values themselves in a \"fallback\" result that is reported in the submission when the fit results are deemed too unreliable (according to a set of selections defined below).\n\nThe initial estimates of the in-transit and out-of-transit regions are also used to compute initial values for the parameters of the FGS fit described below. The distance between the means is taken as an initial guess for T in the fit; the midpoint is taken as an initial guess for Tcenter; and the average of the two widths is the initial guess for tau.  An initial guess for the limb-darkening parameter is derived from the range of normalized light-curve intensities present in the in-transit region.  The background polynomial parameters are initialized such that the constant term is 1 and the linear, quadratic, and cubic terms are all 0.  The initial guess for the fit depth is taken from the Gaussian baseline.  Initial guesses for the AIRS fits are taken from the output of the FGS fit.\n\nTiming estimates of the preprocessing phase of the analysis include all of the above operations and are performed in the Kaggle notebook environment with a P100 gpu, usually running on 50-100 examples from the training set.  There was some variability between one timing run and the next, with this phase of the analysis typically clocking in at anywhere between 12 and 15 seconds per transit on average. \n\n# Parametric fit\n\nThe signal is modeled as a trapezoid, multiplied by a parabola centered at the trapezoid's midpoint.  Its shape is characterized by five parameters:\n * Tcenter, the midpoint of the trapezoid and parabola.\n * T, the duration of the transit, i.e. the time between the middle of the ingress (left trapezoid leg) and the middle of the egress (right trapezoid leg)\n * tau, the ingress/egress time.  This parameter is shared between the ingress and egress regions; that is, the signal model does not allow for the ingress and egress times to be different.\n * LDC, the strength of the parabola which models the effect of limb darkening.  Note that this is not the same as any of the parameters of limb darkening models in the literature; it is more properly understood as an empirical parameter which describes the shape of the top of the trapezoid.\n\nThe values of these parameters in the fit are all limited to fixed ranges by transforming the fit parameters through tanh or sigmoid functions. The trapezoid shape is implemented in a fast, gpu-friendly way using pytorch  (code simplified slightly for readability):\n\n  `trapezoid=torch.clip((torch.arange(0,n_timesteps)-tstart)/tau,0,1)*torch.clip((tend-torch.arange(0,n_timesteps))/tau,0,1)`\n\nwhere `tstart` and `tend` are `Tcenter-T/2` and `Tcenter+T/2`, respectively.  The limb-darkening effect is applied as a multiplicative correction to this:\n\n```\nlimbdark=torch.ones((n_timesteps,))\nlimbdark_x=(torch.arange(0,n_timesteps)-Tcenter)/max(T/2,1)\nlimbdark=torch.clip(limbdark-ldc1*limbdark_x**2,0,1)\nsignal_shape=limbdark*trapezoid\n```\n\nThis four-parameter shape is normalized so that its maximum value is 1 and is then combined with the fourth-order polynomial background with the help of another fit parameter, the transit depth:\n\n`fit_y=fit_y*(1-fit_depth*signal_shape)`\n\nThe transit depth is constrained to be positive by transforming the corresponding fit parameter using a softplus function.  Since the model is implemented in pytorch, it is straightforward to provide a gradient estimate to the minimizer using torch's automatic differentiation.  A chisquare function is defined in the usual way as the sum of `(obs_y-fit_y)**2/y_err**2` over all time and wavelength bins.  The model is first fit to the FGS data using scipy.optimize.minimize() to find the fit parameters that minimize chisquare, allowing all parameters to float freely in the fit.  Up to three calls to minimize() are allowed:  the first uses the default minimizer, BFGS; if that fails (or if the corresponding uncertainty analysis, described below, fails), then a second call is made to the SLSQP minimizer; if that fails, then a third call is made to the Powell minimzer.  Since the FGS data consist of only one wavelength, i.e. only one light curve, this fit does not benefit from gpu parallelization and runs slightly faster on the cpu.\n\nThe results from the FGS fit are then used to seed the fits to the AIRS data.  The model is fit to each of the 282 AIRS wavelengths separately, although these fits are all performed in parallel using torch's Adam minimizer.  The AIRS fits run for a fixed number of iterations, but terminate early if the chisquare function stops decreasing quickly enough.  In order to save time on the AIRS fits, the Tcenter and T parameters are fixed to the values obtained from the FGS fit; however, the tau, LDC, and transit depth parameters (along with all four background polynomial parameters) remain free and independent for each wavelength.\n\nUncertainties for all parameters at each wavelength are estimated as the square root of the diagonal elements of the inverse of the Hessian matrix at the minimum.  Due to time constraints, it is not always possible to find a good enough minimum for this procedure to work; even in the best case, where every wavelength finds the correct minimum and a non-singular Hessian matrix, the computation is time-consuming enough that it creates tension with the runtime limits of the competition.  Thus, the uncertainty calculation is done only for every other wavelength, and missing values (whether from failed minimizations or deliberate skipping) are imputed (or extrapolated) from their nearest neighbors.  In the event that all fits for a given transit fail to produce uncertainty estimates, the analysis moves forward using the average uncertainty values observed in successful fits to the training data.\n\nThe following two images show two of the fitted light curves for a typical transit, planet 189083129.   The first image shows the light curve for the FGS, and the second for one of the light curves from AIRS, wavelength 150.  Both are well-described by the fit.  The grey bands show the boundaries of the initial transit detection.\n\n![Fit to FGS light curve for planet 189083129](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Fc04b787d3c4568c3c463159dc8c3a9d5%2Ffgs_light_curve_fit_planet_1890803129.png?generation=1758850162966272&alt=media)\n\n![Fit to light curve for AIRS wavelength 150 for planet 189083129](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Ffb81891486462a609aaa66a2d654c533%2Fch0_fit_planet1890803129_visit0_wavelength150.png?generation=1758850257917836&alt=media)\n\nThe following two images show the corresponding fits for a pathological case with truncated out-of-transit regions, planet 158006264 visit 1:\n\n![Fit to FGS light curve for planet 158006264](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2F78aff473df8d1c6ce8287291a1b338c6%2Ffgs_light_curve_fit_planet_158006264.png?generation=1758850499483929&alt=media)\n\n![Fit to light curve for AIRS wavelength 150 for planet 158006264](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2F4f6a3b710ef1ca82c373a9240a0ca937%2Fch0_fit_planet158006264_visit1_wavelength150.png?generation=1758850555301923&alt=media)\n\nThere is no special handling in the fitter for these pathological cases, but in the few cases I've looked at, the description of the light curve is nevertheless reasonable if the fitter achieves good convergence. \n\nIn the offline environment used for training, the fitter typically achieves an adequate but not fantastic fit quality, although there are some cases where the fit quality is worse.  The following figure shows the distribution of chisquare/ndof for all planets and all wavelengths in the training set:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F84327%2Fd843d6cc55983f7a2f9e59cf186ffbb4%2Fchisq_ndof.png?generation=1758850810978620&alt=media)\n\nI have not attempted to reproduce this plot in the Kaggle notebook environment.\n\nIn the Kaggle notebook environment, using a P100 gpu, the fit to the FGS data (including its uncertainty analysis) typically takes about 2 or 3 seconds to process one transit's light curve, while the parallel fit to all AIRS light curves typically takes about 6 or 7 seconds.  These figures are averages; there is substantial variability among transits.\n\n#Neural nets for smoothing, limb darkening, and bias correction\n\nAlthough the raw results from the fit include a prediction and uncertainty for each wavelength, they do not leverage the statistical power of the observations at neighboring wavelengths, nor do they account for effects like limb darkening.  Two neural nets are trained to handle these tasks.  Both have roughly the same architecture: they take as input a set of observations for each wavelength; an MLP maps these inputs to a length-64 feature vector for each wavelength, and then a two-layer, 1-dimensional convolutional network predicts a residual correction to the initial spectrum and its uncertainty estimates.  The two networks differ only in their inputs; one constitutes the \"main\" analysis and the other is a fallback in case of fit failure.  The main network takes as input several features for each wavelength:\n  * the transit depth estimate and its uncertainty, both from the Gaussian baseline\n  * the orbital parameters\n  * the mean of the light curve before the light curve is normalized to 1\n  * the best-fit values of the background polynomial parameters and their fit uncertainty estimates (based on the Hessian)\n  * residuals, i.e. `(fit-observation)/uncertainty` in the pre-transit and post-transit regions as well as in a few subsets of the in-transit region (regions re-defined based on the output of the fit)\n  * the average fitted background and signal in these regions (not including signal in the pre- and post-transit regions, since that is by definition zero)\n  * the total transit duration and its uncertainty (based on the FGS fit)\n  * several observables describing the location of the image of the star on the spatial directions of the FGS and AIRS detectors:  mean and width of the bright spot, averaged over time, and the corresponding standard deviations and ranges (maximum across time minus minimum across time), as well as the number of masked pixels in this transit\n  * the fitted values of the transit depth, tau, and LDC parameters, as well as their uncertainties at this wavelength\n  * the fitted values of the transit depth, tau, and LDC parameters and thier uncertainties, averaged over sliding windows of size 3, 7, and 47 on the frequency axis\n\nThe main network predicts a residual correction to the fitted transit depth and its fit uncertainty.\n\nThe fallback network mirrors the main network, but it takes a reduced set of inputs which eliminate any features derived from the fit, so it is left with only the depth estimates and uncertainties from the baseline analysis and the spatial observables; it predicts a residual correction to the Gaussian baseline estimates instead of the values from the fit.\n\nDuring training, the labels are smeared as a means of data augmentation.  Before processing each batch, a random offset is sampled for each wavelength from a Gaussian distribution centered at zero with width given by a multiple of the fit uncertainty.  The multiple is 5 at the beginning of training, and it attenuates by a factor of 0.9998 after each epoch.  Training runs for 150k epochs.  \n\nAt the beginning of training, the predicted spectra often differ from the labels by several times the predicted uncertainty, leading to large gradients.  To stabilize training at the beginning of the run, the uncertainty from the fit (to which the residual correction predicted by the network is added) is artificially increased by a factor which starts at 1000 and decreases linearly to zero over the course of 1000 epochs. \n\nBoth of these networks are trained using a five-fold cross-validation, and for each network, the residual corrections predicted by the five folds are ensembled using a simple arithmetic mean.\n\nThe fallback network is deemed unreliable if the initial-guess widths of the ingress and egress ratio are very different from each other, or if the initial-guess ingress or egress region is too close to the beginning or end of the observation window.   It is deemed reliable otherwise.\n\nThe result reported to the submission.csv for a given wavelength is the one from the main analysis unless the fallback is deemed reliable and any of the following conditions are met:\n  * the sum of the linear, quadratic, and cubic polynomial coefficients for that wavelength is larger than 0.07\n  * the chisquare/ndof for the fit at that wavelength is larger than 1.5\n  * the fitted value of the tau parameter is larger than 40\n  * the fit uncertainty on the tau parameter is larger than 5\n  * any of the residuals input to the main neural net has a magnitude larger than 1\n  * the detected start or end of transit is too close to the edge of the observation window.  (This condition is an oversight; it should have been removed in light of the very similar but not identical reliability cut for the fallback analysis.)\n  * the LDC parameter is less than 0.1\n\nThese selections are applied on a wavelength-by-wavelength basis.  Most of these cuts have only a small effect on the training set; most are motivated by a desire that the main network should not be asked to make a prediction on regions of its input space where there was very little training data.\n\nThe neural net operations described above typically take about 1.5-2 seconds per transit in timing tests in the Kaggle notebook environment with a P100 gpu.\n\n# Combination of multiple visits\n\nIn cases where a planetary system was observed in multiple transits, the combined prediction was an error-weighted average of the predictions from the two transits, taking half the difference between the central values of the two predictions as a systematic uncertainty added in quadrature to the uncertainty obtained from propagation of uncertainty.\n\n# An unexpected snag\n\nValidation checks performed while porting the code to the Kaggle notebook environment turned up a curious difference in behavior for the fitting portion of this analysis:  the minimizers generally did not do as good a job finding minima in the Kaggle environment as they did in the offline environment.  In particular, the chisquare/ndof values from the fits in the Kaggle environment were still generally ok, but they weren't quite as good as the were offline, and the corresponding uncertainty estimates were often a bit bigger.  As this issue only presented itself very close to the close of the competition, I did not manage to pin down the exact reason why.  I wasn't able to spot a difference in the inputs or the libraries that I know were used (I checked pytorch, numpy, scipy, and cuda/cudnn versions, but not other supporting libraries.)  I was able to mitigate this effect a bit by adding checks to the notebook code that would trigger another call to the fitter in order to improve convergence when needed, but I wasn't able to convince myself that the problem was really solved.  I suspect this may be part of why there is a large gap between my local cross-validation scores and the actual leaderboard score that I got in the end.  If anyone ran into similar difficulties, I'd be curious to hear about it.\n\n# Code availability\n\nI have attached to this writeup a link to the submission notebook.  Training code and model weights are available upon request, but I think that the parts anyone might be interested in and able to reuse elsewhere will likely be in the notebook.\n"
  }
}