{
  "id": 157014,
  "title": "QWK not compatible with TPU in Tensorflow",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/157014",
  "author_name": "Pasquale",
  "post_date": "2020-06-08T22:53:30.411000",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>The Kappa metric implemented in Tensorflow addons (<a href=\"https://www.tensorflow.org/addons/api_docs/python/tfa/metrics/CohenKappa\">Here</a>) is currently not supported on TPUs. Does anyone know if there exists a compatible version or how to make one?\nMy understanding is that the one included in the library is not compatible because it makes use of <code>tf.confusion_matrix</code>.\nThanks!</p>",
  "messages": [
    {
      "id": 878841,
      "postDate": "2020-06-08T22:53:30.410Z",
      "content": "<p>The Kappa metric implemented in Tensorflow addons (<a href=\"https://www.tensorflow.org/addons/api_docs/python/tfa/metrics/CohenKappa\">Here</a>) is currently not supported on TPUs. Does anyone know if there exists a compatible version or how to make one?\nMy understanding is that the one included in the library is not compatible because it makes use of <code>tf.confusion_matrix</code>.\nThanks!</p>",
      "rawMarkdown": "The Kappa metric implemented in Tensorflow addons ([Here](https://www.tensorflow.org/addons/api_docs/python/tfa/metrics/CohenKappa)) is currently not supported on TPUs. Does anyone know if there exists a compatible version or how to make one?\nMy understanding is that the one included in the library is not compatible because it makes use of ` tf.confusion_matrix`.\nThanks!",
      "votes": 2
    },
    {
      "id": 878875,
      "postDate": "2020-06-09T01:01:01.897Z",
      "content": "<p>```\ndef cohen_kappa(num_classes=num_classes):\n    def quadratic_kappa_score(y_true, y_pred):\n        eps = 1e-6\n        y_pred = tf.math.round(y_pred)</p>\n\n<pre><code>    weights = tf.range(0, num_classes, dtype=\"float32\") / (num_classes - 1)\n    weights = (weights - tf.expand_dims(weights, -1)) ** 2\n\n    hist_true = tf.math.reduce_sum(y_true, axis=0)\n    hist_pred = tf.math.reduce_sum(y_pred, axis=0)\n\n    E = tf.expand_dims(hist_true, axis=-1) * hist_pred\n    E = E / (tf.math.reduce_sum(E, keepdims=False) + eps)\n\n    O = tf.transpose(tf.transpose(y_true) @ y_pred)\n    O = O / (tf.math.reduce_sum(O) + eps)\n\n    num = weights * O\n    den = weights * E\n\n    QWK = (1 - tf.math.reduce_sum(num) / (tf.math.reduce_sum(den) + eps))\n    return QWK\nreturn quadratic_kappa_score\n</code></pre>\n\n<p>```</p>\n\n<p>This is what I use for one hot data.</p>",
      "rawMarkdown": "```\ndef cohen_kappa(num_classes=num_classes):\n    def quadratic_kappa_score(y_true, y_pred):\n        eps = 1e-6\n        y_pred = tf.math.round(y_pred)\n        \n        weights = tf.range(0, num_classes, dtype=\"float32\") / (num_classes - 1)\n        weights = (weights - tf.expand_dims(weights, -1)) ** 2\n\n        hist_true = tf.math.reduce_sum(y_true, axis=0)\n        hist_pred = tf.math.reduce_sum(y_pred, axis=0)\n\n        E = tf.expand_dims(hist_true, axis=-1) * hist_pred\n        E = E / (tf.math.reduce_sum(E, keepdims=False) + eps)\n\n        O = tf.transpose(tf.transpose(y_true) @ y_pred)\n        O = O / (tf.math.reduce_sum(O) + eps)\n\n        num = weights * O\n        den = weights * E\n\n        QWK = (1 - tf.math.reduce_sum(num) / (tf.math.reduce_sum(den) + eps))\n        return QWK\n    return quadratic_kappa_score\n```\n\nThis is what I use for one hot data.",
      "replies": [
        {
          "id": 878897,
          "postDate": "2020-06-09T01:56:52.293Z",
          "content": "<p>Thank you very much! I'll give it a try and let you know if I manage to make it work :)</p>",
          "rawMarkdown": "Thank you very much! I'll give it a try and let you know if I manage to make it work :)"
        },
        {
          "id": 881162,
          "postDate": "2020-06-10T19:29:56.080Z",
          "content": "<p><a href=\"/richardxiao03\">@richardxiao03</a> Are you sure it computes the QWK correctly? I only added these few lines, since I'm using it with regression:\n<code>\n        y_true = tf.reshape(y_true, [-1])\n        y_pred = tf.reshape(y_pred, [-1])\n        y_pred=tf.clip_by_value(y_pred, 0, 5, name=None)\n        y_true=tf.one_hot(tf.cast(y_true, dtype='int32'), num_classes)\n        y_pred=tf.one_hot(tf.cast(y_pred, dtype='int32'), num_classes)\n</code> \nand it gives a different (lower) QWK compared to the one computed by <code>sklearn.metrics.cohen_kappa_score</code>. It is possible I made a mistake somewhere else though.\nEdit: I indeed made a mistake, I should have rounded before the one_hot function :D</p>",
          "rawMarkdown": "@richardxiao03 Are you sure it computes the QWK correctly? I only added these few lines, since I'm using it with regression:\n``` \n        y_true = tf.reshape(y_true, [-1])\n        y_pred = tf.reshape(y_pred, [-1])\n        y_pred=tf.clip_by_value(y_pred, 0, 5, name=None)\n        y_true=tf.one_hot(tf.cast(y_true, dtype='int32'), num_classes)\n        y_pred=tf.one_hot(tf.cast(y_pred, dtype='int32'), num_classes)\n``` \nand it gives a different (lower) QWK compared to the one computed by ` sklearn.metrics.cohen_kappa_score`. It is possible I made a mistake somewhere else though.\nEdit: I indeed made a mistake, I should have rounded before the one_hot function :D",
          "votes": 2
        },
        {
          "id": 885084,
          "postDate": "2020-06-13T21:27:46.573Z",
          "content": "<p>Actually, I think there is an error with the original implementation. I am pretty sure that the metric is calculated over a batch rather than the entire data set, so the score you see will be an average of the batch scores. However, averaging the scores will lead to different results than calculating the score over the entire data set. For this reason, I updated the original code (which I modified from somewhere else) so that the metric can keep track of its state and give a more accurate score. The code is below.\n```\nfrom sklearn.metrics import cohen_kappa_score</p>\n\n<p>class CohenKappa(tf.keras.callbacks.Callback):\n    '''\n    Computes Cohen's Kappa score with quadratic weighting for one hot encoded data.\n    data is the data set to calculate scores.\n    labels is the set of labels corresponding to the data provided.\n    num_classes is the number of classes in the data.\n    file_path is the path to save model.\n    '''\n    def <strong>init</strong>(self, data, labels, num_classes=num_classes, file_path='model.h5', *<em>kwargs):\n        super(CohenKappa, self).<strong>init</strong>(</em>*kwargs)\n        self.data = data\n        self.file_path = file_path\n        self.num_classes = num_classes\n        self.scores = []\n        self.y_true = labels</p>\n\n<pre><code>def on_train_begin(self, logs=None):\n    self.best = np.NINF\n\ndef on_epoch_end(self, epoch, logs=None):\n    y_pred = self.model.predict(self.data)\n    y_pred = np.argmax(y_pred, -1)\n\n    QWK = cohen_kappa_score(self.y_true, y_pred, weights='quadratic')\n    if np.greater(QWK, self.best):\n        print(f'\\nModel improved from {self.best} to {QWK}. Saving model to {self.file_path}')\n        self.best = QWK\n        self.model.save(self.file_path)\n    else:\n        print(f'\\nModel did not improve from {self.best}. Kappa score {QWK}')\n    self.scores.append(QWK)\n</code></pre>\n\n<p>```\nIf you decide to use this, let me know if you run into any issues!\nEDIT: I just updated this code so that it is now a callback rather than a metric. What I have found is that implementing a custom metric to calculate scores over an entire data set is very convoluted whereas a custom callback is much easier to implement. The metric seems to not work as to my understanding with how state updates should be made. To fix this, I created a callback that will make predictions on the validation set after each epoch. This means that you won't be able to see the score for the training set. This callback will also save the best version of your model.</p>",
          "rawMarkdown": "Actually, I think there is an error with the original implementation. I am pretty sure that the metric is calculated over a batch rather than the entire data set, so the score you see will be an average of the batch scores. However, averaging the scores will lead to different results than calculating the score over the entire data set. For this reason, I updated the original code (which I modified from somewhere else) so that the metric can keep track of its state and give a more accurate score. The code is below.\n```\nfrom sklearn.metrics import cohen_kappa_score\n\nclass CohenKappa(tf.keras.callbacks.Callback):\n    '''\n    Computes Cohen's Kappa score with quadratic weighting for one hot encoded data.\n    data is the data set to calculate scores.\n    labels is the set of labels corresponding to the data provided.\n    num_classes is the number of classes in the data.\n    file_path is the path to save model.\n    '''\n    def __init__(self, data, labels, num_classes=num_classes, file_path='model.h5', **kwargs):\n        super(CohenKappa, self).__init__(**kwargs)\n        self.data = data\n        self.file_path = file_path\n        self.num_classes = num_classes\n        self.scores = []\n        self.y_true = labels\n\n    def on_train_begin(self, logs=None):\n        self.best = np.NINF\n\n    def on_epoch_end(self, epoch, logs=None):\n        y_pred = self.model.predict(self.data)\n        y_pred = np.argmax(y_pred, -1)\n        \n        QWK = cohen_kappa_score(self.y_true, y_pred, weights='quadratic')\n        if np.greater(QWK, self.best):\n            print(f'\\nModel improved from {self.best} to {QWK}. Saving model to {self.file_path}')\n            self.best = QWK\n            self.model.save(self.file_path)\n        else:\n            print(f'\\nModel did not improve from {self.best}. Kappa score {QWK}')\n        self.scores.append(QWK)\n```\nIf you decide to use this, let me know if you run into any issues!\nEDIT: I just updated this code so that it is now a callback rather than a metric. What I have found is that implementing a custom metric to calculate scores over an entire data set is very convoluted whereas a custom callback is much easier to implement. The metric seems to not work as to my understanding with how state updates should be made. To fix this, I created a callback that will make predictions on the validation set after each epoch. This means that you won't be able to see the score for the training set. This callback will also save the best version of your model.",
          "votes": 2
        },
        {
          "id": 885161,
          "postDate": "2020-06-14T01:02:23.733Z",
          "content": "<p>You are right, it didn't compute the right value over batches. I'll try this implementation tomorrow and let you know of any issues I find. Thanks for sharing! :)</p>",
          "rawMarkdown": "You are right, it didn't compute the right value over batches. I'll try this implementation tomorrow and let you know of any issues I find. Thanks for sharing! :)"
        },
        {
          "id": 885187,
          "postDate": "2020-06-14T02:03:51.300Z",
          "content": "<p>I made some adjustments to the code above. You will find most details in the EDIT section. In short, I found that a custom metric is harder to implement than just a callback. Again, let me know if you have trouble implementing it!</p>",
          "rawMarkdown": "I made some adjustments to the code above. You will find most details in the EDIT section. In short, I found that a custom metric is harder to implement than just a callback. Again, let me know if you have trouble implementing it!",
          "votes": 1
        },
        {
          "id": 923469,
          "postDate": "2020-07-10T22:13:08.103Z",
          "content": "<p><a href=\"/richardxiao03\">@richardxiao03</a> I still have a (relatively small) difference between the QWK detected during the training on the validation set and the QWK computed on the validation set <em>outside</em> of the training phase, after loading the weights saved by your callback. I think it could be cause by the different way in which batch normalization and dropout layers work during the training and inference phases. Or maybe I have a bug somewhere :D\nHave you checked if your QWK values are consistent?</p>",
          "rawMarkdown": "@richardxiao03 I still have a (relatively small) difference between the QWK detected during the training on the validation set and the QWK computed on the validation set *outside* of the training phase, after loading the weights saved by your callback. I think it could be cause by the different way in which batch normalization and dropout layers work during the training and inference phases. Or maybe I have a bug somewhere :D\nHave you checked if your QWK values are consistent?"
        },
        {
          "id": 932020,
          "postDate": "2020-07-16T16:20:06.120Z",
          "content": "<p>Well, the scores are computed using sklearn's metric so the difference can't be with the way it is calculated. I am getting consistent scores during and after training. Maybe you are evaluating using different saved models. Perhaps you are using the saved model from the callback during training while you are using the most recent model after training.</p>",
          "rawMarkdown": "Well, the scores are computed using sklearn's metric so the difference can't be with the way it is calculated. I am getting consistent scores during and after training. Maybe you are evaluating using different saved models. Perhaps you are using the saved model from the callback during training while you are using the most recent model after training.",
          "votes": 1
        },
        {
          "id": 932047,
          "postDate": "2020-07-16T16:43:44.670Z",
          "content": "<p>No, the difference is not in the way the QWK is computed but between the predictions made by the network inside the callback and at the end of the training, when I load the best weights and do the inference again on the validation set. I double checked the code and it looks fine, that's why I thought the difference was caused by the different behavior of some layers between the training and inferencing modes. Obviously it's still possible that there is a bug that I haven't spotted :D\nAnyway, thank you so much for your time and for looking into that</p>",
          "rawMarkdown": "No, the difference is not in the way the QWK is computed but between the predictions made by the network inside the callback and at the end of the training, when I load the best weights and do the inference again on the validation set. I double checked the code and it looks fine, that's why I thought the difference was caused by the different behavior of some layers between the training and inferencing modes. Obviously it's still possible that there is a bug that I haven't spotted :D\nAnyway, thank you so much for your time and for looking into that"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 878875,
      "author_name": "Richard Xiao",
      "author_url": "",
      "post_date": "2020-06-09T01:01:01.897000",
      "content": "<p>```\ndef cohen_kappa(num_classes=num_classes):\n    def quadratic_kappa_score(y_true, y_pred):\n        eps = 1e-6\n        y_pred = tf.math.round(y_pred)</p>\n\n<pre><code>    weights = tf.range(0, num_classes, dtype=\"float32\") / (num_classes - 1)\n    weights = (weights - tf.expand_dims(weights, -1)) ** 2\n\n    hist_true = tf.math.reduce_sum(y_true, axis=0)\n    hist_pred = tf.math.reduce_sum(y_pred, axis=0)\n\n    E = tf.expand_dims(hist_true, axis=-1) * hist_pred\n    E = E / (tf.math.reduce_sum(E, keepdims=False) + eps)\n\n    O = tf.transpose(tf.transpose(y_true) @ y_pred)\n    O = O / (tf.math.reduce_sum(O) + eps)\n\n    num = weights * O\n    den = weights * E\n\n    QWK = (1 - tf.math.reduce_sum(num) / (tf.math.reduce_sum(den) + eps))\n    return QWK\nreturn quadratic_kappa_score\n</code></pre>\n\n<p>```</p>\n\n<p>This is what I use for one hot data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 878897,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-09T01:56:52.293000",
          "content": "<p>Thank you very much! I'll give it a try and let you know if I manage to make it work :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 881162,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-10T19:29:56.080000",
          "content": "<p><a href=\"/richardxiao03\">@richardxiao03</a> Are you sure it computes the QWK correctly? I only added these few lines, since I'm using it with regression:\n<code>\n        y_true = tf.reshape(y_true, [-1])\n        y_pred = tf.reshape(y_pred, [-1])\n        y_pred=tf.clip_by_value(y_pred, 0, 5, name=None)\n        y_true=tf.one_hot(tf.cast(y_true, dtype='int32'), num_classes)\n        y_pred=tf.one_hot(tf.cast(y_pred, dtype='int32'), num_classes)\n</code> \nand it gives a different (lower) QWK compared to the one computed by <code>sklearn.metrics.cohen_kappa_score</code>. It is possible I made a mistake somewhere else though.\nEdit: I indeed made a mistake, I should have rounded before the one_hot function :D</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 885084,
          "author_name": "Richard Xiao",
          "author_url": "",
          "post_date": "2020-06-13T21:27:46.573000",
          "content": "<p>Actually, I think there is an error with the original implementation. I am pretty sure that the metric is calculated over a batch rather than the entire data set, so the score you see will be an average of the batch scores. However, averaging the scores will lead to different results than calculating the score over the entire data set. For this reason, I updated the original code (which I modified from somewhere else) so that the metric can keep track of its state and give a more accurate score. The code is below.\n```\nfrom sklearn.metrics import cohen_kappa_score</p>\n\n<p>class CohenKappa(tf.keras.callbacks.Callback):\n    '''\n    Computes Cohen's Kappa score with quadratic weighting for one hot encoded data.\n    data is the data set to calculate scores.\n    labels is the set of labels corresponding to the data provided.\n    num_classes is the number of classes in the data.\n    file_path is the path to save model.\n    '''\n    def <strong>init</strong>(self, data, labels, num_classes=num_classes, file_path='model.h5', *<em>kwargs):\n        super(CohenKappa, self).<strong>init</strong>(</em>*kwargs)\n        self.data = data\n        self.file_path = file_path\n        self.num_classes = num_classes\n        self.scores = []\n        self.y_true = labels</p>\n\n<pre><code>def on_train_begin(self, logs=None):\n    self.best = np.NINF\n\ndef on_epoch_end(self, epoch, logs=None):\n    y_pred = self.model.predict(self.data)\n    y_pred = np.argmax(y_pred, -1)\n\n    QWK = cohen_kappa_score(self.y_true, y_pred, weights='quadratic')\n    if np.greater(QWK, self.best):\n        print(f'\\nModel improved from {self.best} to {QWK}. Saving model to {self.file_path}')\n        self.best = QWK\n        self.model.save(self.file_path)\n    else:\n        print(f'\\nModel did not improve from {self.best}. Kappa score {QWK}')\n    self.scores.append(QWK)\n</code></pre>\n\n<p>```\nIf you decide to use this, let me know if you run into any issues!\nEDIT: I just updated this code so that it is now a callback rather than a metric. What I have found is that implementing a custom metric to calculate scores over an entire data set is very convoluted whereas a custom callback is much easier to implement. The metric seems to not work as to my understanding with how state updates should be made. To fix this, I created a callback that will make predictions on the validation set after each epoch. This means that you won't be able to see the score for the training set. This callback will also save the best version of your model.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 885161,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-06-14T01:02:23.733000",
          "content": "<p>You are right, it didn't compute the right value over batches. I'll try this implementation tomorrow and let you know of any issues I find. Thanks for sharing! :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 885187,
          "author_name": "Richard Xiao",
          "author_url": "",
          "post_date": "2020-06-14T02:03:51.300000",
          "content": "<p>I made some adjustments to the code above. You will find most details in the EDIT section. In short, I found that a custom metric is harder to implement than just a callback. Again, let me know if you have trouble implementing it!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 923469,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-07-10T22:13:08.103000",
          "content": "<p><a href=\"/richardxiao03\">@richardxiao03</a> I still have a (relatively small) difference between the QWK detected during the training on the validation set and the QWK computed on the validation set <em>outside</em> of the training phase, after loading the weights saved by your callback. I think it could be cause by the different way in which batch normalization and dropout layers work during the training and inference phases. Or maybe I have a bug somewhere :D\nHave you checked if your QWK values are consistent?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 932020,
          "author_name": "Richard Xiao",
          "author_url": "",
          "post_date": "2020-07-16T16:20:06.120000",
          "content": "<p>Well, the scores are computed using sklearn's metric so the difference can't be with the way it is calculated. I am getting consistent scores during and after training. Maybe you are evaluating using different saved models. Perhaps you are using the saved model from the callback during training while you are using the most recent model after training.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 932047,
          "author_name": "Pasquale",
          "author_url": "",
          "post_date": "2020-07-16T16:43:44.670000",
          "content": "<p>No, the difference is not in the way the QWK is computed but between the predictions made by the network inside the callback and at the end of the training, when I load the best weights and do the inference again on the validation set. I double checked the code and it looks fine, that's why I thought the difference was caused by the different behavior of some layers between the training and inferencing modes. Obviously it's still possible that there is a bug that I haven't spotted :D\nAnyway, thank you so much for your time and for looking into that</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "878841": "The Kappa metric implemented in Tensorflow addons ([Here](https://www.tensorflow.org/addons/api_docs/python/tfa/metrics/CohenKappa)) is currently not supported on TPUs. Does anyone know if there exists a compatible version or how to make one?\nMy understanding is that the one included in the library is not compatible because it makes use of ` tf.confusion_matrix`.\nThanks!",
    "878875": "```\ndef cohen_kappa(num_classes=num_classes):\n    def quadratic_kappa_score(y_true, y_pred):\n        eps = 1e-6\n        y_pred = tf.math.round(y_pred)\n        \n        weights = tf.range(0, num_classes, dtype=\"float32\") / (num_classes - 1)\n        weights = (weights - tf.expand_dims(weights, -1)) ** 2\n\n        hist_true = tf.math.reduce_sum(y_true, axis=0)\n        hist_pred = tf.math.reduce_sum(y_pred, axis=0)\n\n        E = tf.expand_dims(hist_true, axis=-1) * hist_pred\n        E = E / (tf.math.reduce_sum(E, keepdims=False) + eps)\n\n        O = tf.transpose(tf.transpose(y_true) @ y_pred)\n        O = O / (tf.math.reduce_sum(O) + eps)\n\n        num = weights * O\n        den = weights * E\n\n        QWK = (1 - tf.math.reduce_sum(num) / (tf.math.reduce_sum(den) + eps))\n        return QWK\n    return quadratic_kappa_score\n```\n\nThis is what I use for one hot data."
  }
}