{
  "topic": {
    "id": 672528,
    "title": "[Math Corpus Prize] Annotated Solution Traces",
    "authorName": "Tong Hui Kang",
    "commentCount": 16,
    "votes": 31,
    "postDate": "2026-02-08T23:50:11.751000"
  },
  "comments": [
    {
      "id": 3450911,
      "authorName": "Sam Bealing",
      "votes": 1,
      "postDate": "2026-04-30T16:11:44.280000",
      "content": "<p>Following up on <a href=\"https://www.kaggle.com/philipvonderlind\" target=\"_blank\">@philipvonderlind</a> questions, I have some further questions to help us assess your dataset more closely, as described in the updated <a href=\"https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3/discussion/695434\" target=\"_blank\">discussion post</a>.</p>\n<ol>\n<li><p>You mention that you add <code>mod 99991</code> to PE problems 'as long as it does not confuse the answer requirement'. In most cases where it is not applied, it is because there is already one <code>mod</code> request but there are other cases where it is not applied (e.g. <code>pe-477</code> and <code>pe-492</code>). Was there any reason for this decision?</p></li>\n<li><p><code>For AIMO3 and IMO problems, I change numerical variables to create parametric variants.</code> - Could you give more detail on how you went about this process (e.g. Why did you select the problems you did? How did you decide which parameter(s) to tweak and to what values?). Did you find any interesting results about how the models solved these variants (it would be particularly interesting if a model found certain parameters significantly harder than others). It would also be interesting to see the code used to generate these problems so this could be used to test even more variants and/or as a template for others looking to run parametrised problems.</p></li>\n<li><p>You demonstrate that your finetuned model has improved performance on the reference questions 9 and 10 as well as the public leaderboard. Did you try running on any other datasets where there is no risk of contamination (as there potentially is for the reference problems even if they exact problems don't appear in the training) and where you can run a larger number of times to reduce effects of any random variance (we have seen large variance in public leaderboard scores and so while the score differences are fairly convincing, with only 2/3 runs there could be some randomness at play)?</p></li>\n</ol>\n<p>Look forward to reading your response on these.</p>"
    },
    {
      "id": 3451206,
      "authorName": "Tong Hui Kang",
      "votes": 0,
      "postDate": "2026-05-01T02:17:05.657000",
      "content": "<h4>1) Modulo 99991</h4>\n<p>Thanks for the sharp observations. My process when applying modulo is a heuristic and I err on not applying the modulo (if there is a mention of modulo anywhere in the question I do not apply the modulo).</p>\n<p>There is the question of why I even applied modulo 99991 which I did not initially elaborate. One reason is to train the model to follow the modulo instruction, another reason is to allow answers to be verified without publishing the answer (for reasons you can see in this thread).</p>\n<h4>2) Parametric variants</h4>\n<p>There are only a few problems that made the parametric variants, which I hand selected. The parametric variants were instructed to be generated with AI coding tools. There are not many recent IMO problems that can be readily transformed into AIMO problem format and not easily solvable by gpt-oss-120b.</p>\n<p>These are a few observations that you might find interesting.</p>\n<p>For reference problem 9, it seems that gpt-oss-120b could perform perfectly for m=0 and m=1 because it can run a brute force approach. The full problem asks for m=8. Interestingly, the performance for m=50 might be better than m=10 because the agent is no longer able to hand count for the answer.</p>\n<p>For the famous IMO 2025 p6, you see that gpt-oss-120b solves n=4 reliably, but they almost always fail at n=9, and the full problem asks for n=2025.</p>\n<p>You can observe that I try to create parametric variants where gpt-oss-120b could solve around half the time.</p>\n<p>I think there is great value in creating parametric variants. This is similar to how you teach humans: you get them to solve a small but still challenging instance of the problem, help them forget about the problem but learn the approach (clear context and update weights), and get them to solve a slightly more challenging instance, and eventually they learn to solve the full problem and hopefully similar problems. I should have spent my December holidays differently by creating these parametric variants for all Euler problems.</p>\n<h4>3) Performance</h4>\n<p>I did show an improvement on reference questions 9 and 10 even though the problem it trained on are parametric variants. I did not claim my model improved public leaderboard performance (I said that it does not degrade it).</p>\n<p>I do not think my model should be performing better than gpt-oss-120b in the public leaderboard. The <a href=\"https://aimo.huikang.dev/corpus.html\" target=\"_blank\">training data</a> only contains the reference problems.</p>\n<p>I have not run my model on other datasets because I predicted that it would not be better either. I wanted to achieve the ability to improve on O(100) problems before I consider an uncontaminated evaluation.</p>\n<p>I did try to train the model on more problems. However, as I add more problems in the dataset, the model no longer perform better on reference problems even though they are still in the training set. It might be that I actually need a lot more completions and labels.</p>\n<p>I was betting that with a few good labels the models can efficiently learn and achieve great performance on the problems it has trained on. While 600 entries and 100k completion tokens (on 9 million prompt tokens) is able to improve performance on the reference problems, it could not achieve the near perfect performance I was hoping for.</p>"
    },
    {
      "id": 3454573,
      "authorName": "Sam Bealing",
      "votes": 1,
      "postDate": "2026-05-07T13:28:08.070000",
      "content": "<p>Thanks for the response here. Some interesting observations, particularly around how performance scales (or perhaps doesn't) with the volume of data. I agree it's hard to find problems with parametric variants but it is an interesting area to explore (and one we are also exploring). I wonder if sometimes with final-answer competitions, models could be trained to pattern spot however this becomes a lot trickier when either the pattern doesn't have a nice closed formula (as is the case in the Reference problem you mention) or where the pattern only holds for a certain type of values (square numbers in the case of the IMO problem you reference). For the latter case, unless you hypothesise that is why 2025 has been chosen (which isn't obvious without knowing the solution given 2025 is just the year which is commonly used as a placeholder figure), you wouldn't get the right pattern.</p>"
    },
    {
      "id": 3450207,
      "authorName": "Philip Vonderlind",
      "votes": 1,
      "postDate": "2026-04-29T13:55:31",
      "content": "<p>Hi, thank you for your submission to the Math Corpus Prize. Me and <a href=\"https://www.kaggle.com/sambealing\" target=\"_blank\">@sambealing</a> will post some questions below that help us in assessing your dataset more closely, as described in the <a href=\"https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3/discussion/695434\" target=\"_blank\">updated discussion post.</a></p>\n<ol>\n<li><p>You mention selecting problems from multiple sources. How did you decide which problems to select for your dataset (filtering, similarity matching, deduplication, …), and how did you verify the validity of the labels? Furthermore, did you scrape these problems yourself, or were the problems curated from already existing datasets?</p></li>\n<li><p>For the reflections &amp; annotations, you mention the use of AI-Agents to assist you with the annotation process. How did you verify (or test for a subset) that the generated reflections and annotation comparisons don't make mistakes themselves, as they are also an LLM-based system that might not be able to solve certain problems, but are confident in their ability to do so, even with a wrong outcome?</p></li>\n</ol>\n<p>Looking forward to your response :)</p>"
    },
    {
      "id": 3450272,
      "authorName": "Tong Hui Kang",
      "votes": 0,
      "postDate": "2026-04-29T16:00:11.673000",
      "content": "<h4>1) Problem curation</h4>\n<p>Most problems were created from existing datasets. I did not validate the correctness of the answer because the problems are already reputable competition problems.</p>\n<p>There are some problems that are derived from current problems (AIMO questions 9 and 10, some IMO problems) by changing the variables. The answers are checked with frontier LLMs, with and without access to the editorial.</p>\n<h4>2) Verification</h4>\n<p>Great question.</p>\n<p>I initially tried rolling out from the preferred and dispreferred paragraph to compare whether the eventual answer is correct, but the difference was not immediately significant.</p>\n<p>The difference between labelling and solving from scratch</p>\n<ul>\n<li>the labelling model is more intelligent (frontier models in frontier AI coding tools)</li>\n<li>the model has access to the correct answer and the editorial</li>\n<li>the model has access to other completions</li>\n</ul>\n<p>The idea is instead of training on the whole completion which includes good actions and bad actions, I want to train on only the good actions. It should be an improvement if the actions I select are significantly better than the actions I do not select. I do not need the labels to be flawless.</p>\n<p>Overall, after training, the model performed better on very similar problems it trained on. I should have also trained a version where I train on the full sequence, to compare the approaches.</p>\n<p>After I published the dataset, I tried an approach where the labeling happens when the student model solves the problem. This approach allowed me to significantly increase the correctness rate, and reduce the average length of the correct sequences.</p>"
    },
    {
      "id": 3403824,
      "authorName": "cm391",
      "votes": 2,
      "postDate": "2026-02-09T10:53:55.440000",
      "content": "<blockquote>\n  <p>I release all contents of my submission for the Math Corpus Prize under Apache 2.0.</p>\n</blockquote>\n<p>you cannot just say this… Project Euler is Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0). I hope the hosts will check this for all submissions.</p>\n<p>Please update the licence. Project Euler is a huge effort to all participants and it is unfair to not respect their <a href=\"https://projecteuler.net/copyright\" target=\"_blank\">licence</a>.</p>"
    },
    {
      "id": 3403996,
      "authorName": "Simon Frieder",
      "votes": 2,
      "postDate": "2026-02-09T17:28:02.607000",
      "content": "<p>Thanks for pointing this out! We'll follow up on this and I agree that the OP should fix this!</p>\n<p>I do want to copy two things from <a href=\"https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3/discussion/671893:\" target=\"_blank\">https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3/discussion/671893:</a> </p>\n<ul>\n<li><em>\"we will announce the winner likely quite some time before the end of the competition. The exact timing depends on when all questions we may have will have been resolved. Similarly to a review-rebuttal period (if you have ever submitted to ICML/ICLR/NeurIPS), you can use this time to polish your Kaggle Discussion write-up, as you address the questions we or the community may have, and make it better.\"</em>  </li>\n<li><em>I also want to encourage the community to actively engaged with the candidate submissions for the Math Corpus Prize, to see if those help with their submissions, if they spot errors, etc.</em> as there are limits to how much we can check and catch, but if the community flags thing we will definitely follow up on that in addition to our own checks.</li>\n</ul>\n<p>This being said, I also want to say two other things:</p>\n<ul>\n<li>in this case, I think this error is an easy fix. One merely needs to change the licences at a per-datapoint level and not release the entire dataset with licence X but point out which datapoints are under licence X_1, which under X_2, and so on. </li>\n<li>Generally, I think this is will be problem with many submission that there will be a \"rebranding\" of licences where the dataset is released under licence X that is not compatible with the original licence Y (sometimes compatability is given, but this is a very finicky legal issue). This also affects big major datasets (not naming any here though), so this is something that I won't judge too harshly, and mainly keep in line with common practice in ML. <strong>If there are really problematic incompatibilities, these are major issues --such as the original datapoint being released under a super restrictive licences that allows nothing, and then being re-released for the Math Corpus Prize under a very permissive one such as the MIT licence-- but otherwise I'd hope that the community won't start becoming \"licence police\", as this won't really contribute to what we all care about: Getting better LLMs that do better on math. Other errors, such as wrong claims that cannot be reproduced, missing arguments/graphs, errors in data where the solution doesn't match the problem, are what I think will really help improve the quality of the dataset and the Kaggle Discussion writeup, and what therefore the community should be on the lookout for.</strong> I think we have to treat licencing as a best-effort but accept that we may not always get things right; if someone finds errors we should strive to correct, but we shouldn't spend excessive effort on this. I have added this guidance to <a href=\"https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3/discussion/671893\" target=\"_blank\">https://www.kaggle.com/competitions/ai-mathematical-olympiad-progress-prize-3/discussion/671893</a></li>\n</ul>"
    },
    {
      "id": 3404037,
      "authorName": "Tong Hui Kang",
      "votes": 0,
      "postDate": "2026-02-09T18:58:15.813000",
      "content": "<p>Thanks for chiming in.</p>\n<p>My intention behind my original statement is that I do not intend to gatekeep anyone from using my contributions here to do anything in any way. I have updated the statements, would like advice on how to improve these statements further!</p>"
    },
    {
      "id": 3406042,
      "authorName": "Jasper Dekoninck",
      "votes": 2,
      "postDate": "2026-02-14T12:34:04.817000",
      "content": "<p>Just saw this dataset. While a think it's a great contribution by <a href=\"https://www.kaggle.com/huikang\" target=\"_blank\">@huikang</a>, it does directly violate the competition rules of Project Euler, which has a clear ban on the publication of answers and solutions beyond problem 100.</p>\n<p>We have worked extensively together with the PE community for our <a href=\"https://matharena.ai/?view=problem&amp;comp=euler--euler\" target=\"_blank\">PE benchmark on MathArena</a>. They are really strict on this requirement. For instance, we do not publish any model outputs (not in dataset form, not on our website) and when we published <a href=\"https://matharena.ai/euler/\" target=\"_blank\">our blog post</a>, we had several back and forth with them to ensure the qualitative analysis did not contain any hint for any of the problems.</p>\n<p>The rule was designed to prevent others from having the fun of solving the problems themselves. The dataset here has a different goal, as it is not to inform others. However, given that they also stress in our case to prevent the publication of any solutions or answers, it seems reasonable to surmise that the same restrictions should apply here. It would be quite unfortunate to see our effort go to waste by other datasets that publish all solutions publicly.</p>"
    },
    {
      "id": 3406184,
      "authorName": "Tong Hui Kang",
      "votes": 0,
      "postDate": "2026-02-14T23:21:19.180000",
      "content": "<p>Thanks for reaching out.</p>\n<p>On this matter, I am happy to follow what the competition organizer suggests.</p>\n<p>On my <a href=\"https://aimo.huikang.dev/annotations.html?problem=pe-461&amp;state=eyayt0&amp;action=6tognb\" target=\"_blank\">website</a> and <a href=\"https://github.com/tonghuikang/aimo3\" target=\"_blank\">Github</a>, I have redacted Project Euler related traces and answers, leaving only the problem statement, the correctness judgment, and the token length taken.</p>\n<p>My contribution here is to show how I annotate the traces and train the model, rather than to upload the problems and their solution traces. I think I have enough examples in the AIMO3 problems to demonstrate how I annotate and train.</p>\n<p>The model I shared is also trained <a href=\"https://aimo.huikang.dev/corpus.html?included=true\" target=\"_blank\">mostly</a> on AIMO3 annotations.</p>"
    },
    {
      "id": 3406305,
      "authorName": "Jasper Dekoninck",
      "votes": 0,
      "postDate": "2026-02-15T10:33:32.270000",
      "content": "<p>Amazing! Thanks :)</p>"
    },
    {
      "id": 3406731,
      "authorName": "Simon Frieder",
      "votes": 1,
      "postDate": "2026-02-16T14:07:39.063000",
      "content": "<p>I think this ultimately on whether sharing the answer to the PE is a request rather than a licence-related requirement.</p>\n<ul>\n<li><p>In the former case, it is up to you whether you want to share the answer or not. In the age of LLMs, I'm not sure how realistic it is to keep the internet free of answers, it seems a requirement that incompatible with the pace at which LLMs proceed and the open nature of the internet; at the same time, I'm sympathetic to PE and their intentions of preserving the competition.\nI defer to the <a href=\"https://www.kaggle.com/huikang\" target=\"_blank\">@huikang</a> in this case in how you want to proceed.</p></li>\n<li><p>In the latter case, if, say using/including PE problems mandate you don't publish their answer, then whatever licence requirements come from integrating PE-data in your submission need be respected. (And you need to clarify which datapoints use which license.)</p></li>\n</ul>"
    },
    {
      "id": 3405122,
      "authorName": "wenliangtlh",
      "votes": -2,
      "postDate": "2026-02-12T06:50:21.847000",
      "content": "<p>Hello, could you please explain how you used <code>corpus.csv</code> for training? Did you take the <code>prompt</code> as the input and <code>completion</code> as the output? Did you apply <strong>loss masking</strong> to the prompt? In other words, are <code>&lt;prompt, completion&gt;</code> pairs used as the training data?</p>"
    },
    {
      "id": 3405129,
      "authorName": "Tong Hui Kang",
      "votes": -1,
      "postDate": "2026-02-12T07:18:45.547000",
      "content": "<p>Thank you for your interest!</p>\n<p>Yes, the training data is exactly the prompt completion pairs (also only those with <code>included=True</code>). Loss should not be computed on the prompt. You can see that the <code>num_loss_tokens</code> is much smaller than <code>num_tokens</code> in <a href=\"https://github.com/tonghuikang/aimo3/blob/master/training/sft/default/metrics.jsonl\" target=\"_blank\">metrics.jsonl</a></p>"
    },
    {
      "id": 3405142,
      "authorName": "wenliangtlh",
      "votes": 0,
      "postDate": "2026-02-12T08:07:37.767000",
      "content": "<p>Okay, thanks a lot. I have one question: if we only train the model using  pairs, could this lead to incomplete responses from the model? Because your completion is only a single step, not the full sequence of steps that should come after the prompt.</p>"
    },
    {
      "id": 3405161,
      "authorName": "Tong Hui Kang",
      "votes": -1,
      "postDate": "2026-02-12T09:11:42.317000",
      "content": "<p>This is an interesting question. Notice that my completion does not include the end of text token. This means I do not take a position on what happens after the completion that I am training.</p>\n<p>If you consider a finetuning dataset where every time it is about write a wrong answer, you train the model to continue thinking about the answer. This might improve the accuracy of the model, but this also makes the model much less likely to commit to any answer.</p>"
    }
  ],
  "index": {
    "id": "672528",
    "title": "[Math Corpus Prize] Annotated Solution Traces",
    "authorName": "Tong Hui Kang",
    "commentCount": "16",
    "votes": "31",
    "postDate": "2026-02-08 23:50:11.751000"
  }
}