{
  "competition": "cafa-5-protein-function-prediction",
  "topic_id": "405090",
  "comments": [],
  "messages": [],
  "raw_show": {
    "topic": {
      "id": 405090,
      "title": "Q: information accretion values in IA.txt",
      "authorName": "Igor Pechersky",
      "commentCount": 2,
      "votes": 12,
      "postDate": "2023-04-26T09:32:54.304000"
    },
    "comments": [
      {
        "id": 2236230,
        "authorName": "Clara De Paolis",
        "votes": 5,
        "postDate": "2023-04-26T16:30:23.163000",
        "content": "<blockquote>\n  <p>Terms deep in the ontology appear less frequently, are harder to predict, and thus their weights are larger</p>\n</blockquote>\n<p>This statement is usually true but not always. Thanks for catching this and we will fix the language to make it clearer in the Evaluation section. </p>\n<p>To explain the exceptions to this rule, let's consider the IA calculation that boils down to </p>\n<p>$$ia(X) = \\log_2\\left(\\frac{\\text{number of proteins with parent term(s) of } X}{\\text{number of proteins with term } X}\\right)$$</p>\n<p>Usually terms deeper in the ontology are more rarely seen because they correspond to functions that are more specific and may be harder to determine than those closer to ontology roots. However, sometimes there are terms that are always seen annotated in proteins whenever the parents terms are seen annotated. In those cases number of proteins with parent term(s) of  X is equal to number of proteins with term X and IA is 0. This is the case with GO:0019492 in your example. This term has two parents: GO:0043102 and GO:0006561. There is one protein annotated with these two parent terms (A0A031WDE4). This protein also is annotated with the child term GO:0019492. Thus the counts in the equation are equal and IA=0</p>\n<p>As you point out one of those parent terms (GO:0006561) has a larger IA. This is because we see term GO:0006561 in 24 proteins. The term has 2 parents (GO:0009084 and GO:0006560). We see 26 proteins annotated with both these terms. log2(26+1/24+1) =0.111.  Note the +1 is to regularize the empirical counts.</p>\n<p>The term GO:1901566 is not a parent of GO:0006560, but it is an ancestor of your leaf term GO:0019492 (just through a different path in the ontology: GO:0019492 &gt;GO:0043102&gt;GO:0008652&gt;GO:1901566). So your point stands anyway. I hope the above explanation helps. </p>"
      },
      {
        "id": 2236254,
        "authorName": "Predrag Radivojac",
        "votes": 4,
        "postDate": "2023-04-26T16:53:58.173000",
        "content": "<p>This is an excellent question and an excellent post. One of the reasons this ia(node) term was called information accretion (and not say information content) was that this is the amount of information a term provides beyond what all of its ancestors collectively do (it's sufficient to just look at immediate parents due to annotation propagation to the root). If a term is always present when all its ancestors are present, then its contribution is 0; it does not bring anything to the table since it's always predicted to be present when the parents are present. If not, then term's contribution is above 0.</p>"
      }
    ]
  },
  "topic": null,
  "index": {
    "id": "405090",
    "title": "Q: information accretion values in IA.txt",
    "authorName": "",
    "commentCount": "2",
    "votes": "12",
    "postDate": "2023-04-26 09:32:54.304000"
  }
}