{
  "topic": {
    "id": 644800,
    "title": "Possible evaluation problem",
    "authorName": "Btbpanda",
    "commentCount": 5,
    "votes": 15,
    "postDate": "2025-11-29T23:14:14.339000"
  },
  "comments": [
    {
      "id": 3362938,
      "authorName": "Clara De Paolis",
      "votes": 1,
      "postDate": "2025-12-05T16:07:03.193000",
      "content": "<p>The fix to remove duplicate adjacent nodes has been integrated into the github repo. Thanks for catching this, <a href=\"https://www.kaggle.com/btbpanda\" target=\"_blank\">@btbpanda</a>! </p>"
    },
    {
      "id": 3360961,
      "authorName": "An Phan",
      "votes": 2,
      "postDate": "2025-12-02T19:37:12.620000",
      "content": "<p>The problem occurred for 4 parent-child pairs because these child terms have both \"is_a\" and \"part_of\" relationship with their parent, leading to duplicated children in the <code>children</code> list for the parent term. During topological sort, the parent term reached <code>in_degree</code> of 0 earlier than expected, resulting in it being added to <code>order</code> before exhausting all of its children (and thus smaller index than its children). </p>\n<p>Example: GO:1990700 has both \"is_a\" and \"part_of\" with parent GO:0006325</p>\n<pre><code>[Term]\n: GO:\nname: nucleolar chromatin organization\nnamespace: biological_process\n:  [PMID:]\nsynonym:  EXACT []\nsynonym:  EXACT []\nis_a: GO:0006325 ! chromatin organization\nrelationship: part_of GO:0006325 ! chromatin organization\nrelationship: part_of GO:0007000 ! nucleolus organization\n</code></pre>\n<p>I have tested a temporary fix for the \"Graph\" class, as detailed in <a href=\"https://www.kaggle.com/code/ahphan/metric-implementation-issue-discussion\" target=\"_blank\">my notebook</a>. The <code>children</code> and <code>adj</code> list of each term needs to be de-duplicated after constructing adjacency matrix, before topological sorting. With the topological sort fixed, the remaining propagation process should behave correctly without modifying. We will keep testing this and eventually update the cafaevaluator code to handle this bug.</p>"
    },
    {
      "id": 3360989,
      "authorName": "Clara De Paolis",
      "votes": 1,
      "postDate": "2025-12-02T20:27:19.993000",
      "content": "<p>Interesting edge case, nice find!  I encourage PRs to the repo for any issues found (and fixed).  </p>"
    },
    {
      "id": 3360141,
      "authorName": "An Phan",
      "votes": 2,
      "postDate": "2025-12-01T22:15:12.887000",
      "content": "<p>Hello, thanks for pointing it out. I did a quick check at the code, and I did observe that some pairs of parent-child terms are not indexed like we would expect (child terms indices should be smaller than parent terms). This is probably due to the structure of the GO DAG where a child term sometimes has a parent that are a few levels further up the DAG (but still considered direct parent due to a direct relationship), which I think led to the mix-up in the resulting topological sort (not sure, I'm checking to see if this was the cause).</p>\n<p>I created a toy prediction file which only has one prediction for term <code>GO:0006338</code> (at index 2116 in <code>G.terms_list</code>, and order 25602 in <code>G.order</code>), and the expected result (after propagating) should have its parent term <code>GO:0006325</code> (at index 2112 in <code>G.terms_list</code>, and order 20878 in <code>G.order</code>) with the same score. During propagation, the parent term <code>GO:0006325</code> at order 20878 was completely removed from <code>order_</code> in the \"remove leaves\" step, leading to incorrect propagation. More details in <a href=\"https://www.kaggle.com/code/ahphan/metric-implementation-issue-discussion\" target=\"_blank\">my notebook</a>.</p>\n<p>The fastest way to fix this is just to comment out the \"remove leaves\" part in the function <code>propagate</code> and only loop through indexes in <code>G.order</code>, so no parent terms get removed by accident; however, this fix will slow down the evaluation code. Right now, I am figuring out a way to fix the topological sort so that all child terms get smaller indexes than their parents in <code>G.order</code>. It only affects a very small number of terms, and these terms might not show up in the ground truth leaderboard, so no rescoring for now until I can nail down the cause and a good solution.</p>"
    },
    {
      "id": 3360501,
      "authorName": "Btbpanda",
      "votes": 7,
      "postDate": "2025-12-02T08:56:44.560000",
      "content": "<p>Hi An Phan. </p>\n<p>Thanks for your reply. I checked your quick fix solution, and, unfortunately, can say it is not enough. Even if we don't remove any terms from <code>order_</code>, you can imagine the case:</p>\n<p>There is a graph <code>G</code> with <code>terms_list</code> <code>['a', 'b', 'c']</code>. There is a chain of terms (from parent to child) <code>a -&gt; b -&gt; c</code>, </p>\n<p>There is a ground truth: <code>{a: 0, b: 0, c: 1}</code></p>\n<p>There is G.order, which is wrong: <code>[0, 2, 1]</code>. The correct order supposed to be <code>[2, 1, 0]</code></p>\n<p>Imagine, we applied your fix, and iterate through that order while propagating the ground truth:</p>\n<h3>Iteration 0:</h3>\n<p>We take <code>a</code>  as a parent and assign it with the max children value, which is 0 comes from <code>b</code>. The result will be  <code>{a: 0, b: 0, c: 1}</code>. </p>\n<h3>Iteration 1:</h3>\n<p>We take <code>c</code> as a parent, but there is no children to propagate, so we keep the same result: <code>{a: 0, b: 0, c: 1}</code></p>\n<h3>Iteration 2:</h3>\n<p>We take 'b' as a parent and assign it with the max children value, which is 1 comes from <code>c</code>. The result will be   <code>{a: 0, b: 1, c: 1}</code></p>\n<p>After the final iteration, you can see that result is different from expected which is <code>{a: 1, b: 1, c: 1}</code></p>\n<h1>Hot fix</h1>\n<p>The most simple, however, a bit inefficient solution that works here (only for the given Graph, not general case) is to apply your fix + add propagation twice, so just replace line 152 <code>CAFA-evaluator-PK/src/cafaeval/parser.py</code> together with comment remove leaves part</p>\n<p>with </p>\n<pre><code>propagate(matrix, ontologies[ns], ontologies[ns].order, mode=)\npropagate(matrix, ontologies[ns], ontologies[ns].order, mode=)\nlogging.debug(.(ns, matrix))\n</code></pre>\n<p>To compensate the inefficiency, I suggest you to replace your propagate function with something more efficient, like I did in my CAFA5 code. To be honest, many things could be done more efficient in that code, but let's focus on propagation now😀</p>\n<pre><code> numba  njit, prange\n\n\n ():\n     i  prange(mat.shape[]):\n         mat[i, k] == :\n            \n\n         j  adj:\n             mat[i, j] == :\n                mat[i, k] = \n                \n\n    \n\n\n ():\n     i  prange(mat.shape[]):\n         j  adj:\n             mat[i, j] &gt; mat[i, k]:\n                mat[i, k] = mat[i, j] \n\n    \n\n ():\n\n    fn = prop_max_cpu_binary  binary  prop_max_cpu\n\n     f  G.order:\n        adj = G.terms_list[f][]\n\n         (adj) == :\n            \n\n        fn(mat, f, np.asarray(adj))\n\n    \n</code></pre>\n<p>and then replace the target propagation at line 152 <code>CAFA-evaluator-PK/src/cafaeval/parser.py</code>  with </p>\n<pre><code>propagate_max(matrix, ontologies[ns], binary=) \npropagate_max(matrix, ontologies[ns], binary=) \nlogging.debug(.(ns, matrix))\n</code></pre>\n<p>and replace prediction propagation at line 223 <code>CAFA-evaluator-PK/src/cafaeval/parser.py</code> with</p>\n<pre><code>propagate_max(matrix[ns], ontologies[ns], binary=) \npropagate_max(matrix[ns], ontologies[ns], binary=) \nlogging.debug(.(ns, matrix))\n</code></pre>\n<p>But of course, the most proper way for the research purposes is to fix <code>.top_sort</code> method of <code>Graph</code></p>"
    }
  ],
  "index": {
    "id": "644800",
    "title": "Possible evaluation problem",
    "authorName": "Btbpanda",
    "commentCount": "5",
    "votes": "15",
    "postDate": "2025-11-29 23:14:14.339000"
  }
}