{"cells":[{"metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:38:04.461975Z","start_time":"2020-04-22T04:38:04.456987Z"},"trusted":true},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import cohen_kappa_score, make_scorer\nimport itertools\nimport sklearn\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport random\n\nfrom IPython.core.interactiveshell import InteractiveShell\nInteractiveShell.ast_node_interactivity = \"all\"","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":" ## Table of Contents\n1. [Intuition of QWK](#intuition)\n2. [Step 1: Create the NxN histogram matrix O](#confusion) <br>\n3. [Step 2: Create the Weighted Matrix w](#weighted)<br>\n4. [Step 3: Create the Expected Matrix](#expected) <br>\n5. [Step 4: Final Step: Weighted Kappa formula and Its python codes](#qwk) \n\n    "},{"metadata":{},"cell_type":"markdown","source":"## **Intuition of QWK** <a id=\"intuition\"></a> <a id=\"intuition\"></a>\n\nTLDR: one can skip to the last section on the python code implementation of QWK and also take reference to [CPMP's Fast QWK Computation](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145105) as well.\n\nKappa or Cohen’s Kappa is like classification accuracy, except that it is normalized at the baseline of random chance on your dataset: It basically tells you how much better your classifier is performing over the performance of a classifier that simply guesses at random according to the frequency of each class."},{"metadata":{},"cell_type":"markdown","source":"First off, we define the formula exactly as mentioned. Quoting from the evaluation page: **Submissions are scored based on the quadratic weighted kappa, which measures the agreement between two outcomes. This metric typically varies from 0 (random agreement) to 1 (complete agreement). In the event that there is less agreement than expected by chance, the metric may go below 0.** \n\n**The quadratic weighted kappa is calculated as follows. First, an N x N histogram matrix O is constructed, such that $O_{i,j}$ corresponds to the number of `isup_grade`'s i (actual) that received a predicted value j. An N-by-N matrix of weights, w, is calculated based on the difference between actual and predicted values:**\n\n$$w_{i,j} = \\dfrac{(i-j)^2}{(N-1)^2}$$\n\n<br>\n\n**An N-by-N histogram matrix of expected outcomes, E, is calculated assuming that there is** ***no correlation*** **between values.This is calculated as the outer product between the actual histogram vector of outcomes and the predicted histogram vector, normalized such that E and O have the same sum. **\n\n<br>\n\nFinally, from these three matrices, the quadratic weighted kappa is calculated as:\n\n\n$$\\kappa = 1 - \\dfrac{\\sum_{i,j}\\text{w}_{i,j}O_{i,j}}{\\sum_{i,j}\\text{w}_{i,j}E_{i,j}}$$\n\nwhere $w$ is the weighted matrix, $O$ is the histogram matrix and $E$ being the expected matrix."},{"metadata":{},"cell_type":"markdown","source":"## **Step 1: Create the NxN histogram matrix O** <a id=\"confusion\"></a>"},{"metadata":{},"cell_type":"markdown","source":"***Warning: This is a counter example to the wrong usage of Quadratic Weighted Kappa. We shall see why soon.***\n\n***Warning: This is a counter example to the wrong usage of Quadratic Weighted Kappa. We shall see why soon.***\n\n***Warning: This is a counter example to the wrong usage of Quadratic Weighted Kappa. We shall see why soon.***\n\n<br>\n\n***Reminder: Although it is a counter example, it still illustrates what a NxN histogram matrix is! We will now call our histogram matrix C instead because in actual fact, the histogram matrix is merely a multi class confusion matrix between actual and predicted values***"},{"metadata":{},"cell_type":"markdown","source":"We use a naive example where there are 5 classes (note our competition is_up grade has 6 classes; but this is just an example. Our `y_true` is the ground truth labels and correspondingly, our `y_pred` is the predicted values."},{"metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:15:51.378798Z","start_time":"2020-04-22T04:15:51.372814Z"},"trusted":true},"cell_type":"code","source":"y_true = pd.Series(['cat',  'cat', 'dog', 'cat',   'cat',  'cat', 'pig',  'pig', 'hen', 'pig'], name = 'Actual')\ny_pred   = pd.Series(['bird', 'hen', 'pig','bird',  'bird', 'bird', 'pig', 'pig', 'hen', 'pig'], name = 'Predicted')\n\nprint(\"Ground truth:\\n{}\".format(y_true))\nprint(\"-\"*40)\nprint(\"Predicted Values:\\n{}\".format(y_pred))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"First, an N x N confusion matrix **C** is constructed, such that $\\text{C}_{i,j}$ is the entry that corresponds to the **number of animal i (actual) that received a predicted value j**. "},{"metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:17:36.46162Z","start_time":"2020-04-22T04:17:36.288531Z"},"trusted":true},"cell_type":"code","source":"classes= ['bird','cat','dog','hen', 'pig']\n\n# thank you https://datascience.stackexchange.com/questions/40067/confusion-matrix-three-classes-python \ndef plot_confusion_matrix(cm, classes,\n                          normalize=False,\n                          title='Confusion matrix',\n                          cmap=plt.cm.Blues):\n    \"\"\"\n    This function prints and plots the confusion matrix.\n    Normalization can be applied by setting `normalize=True`.\n    \"\"\"\n    \n    if normalize:\n        cm = cm.astype('float') / cm.sum(axis=1)[:, np.newaxis]\n        print(\"Normalized confusion matrix\")\n    else:\n        print('Confusion matrix, without normalization')\n\n    plt.imshow(cm, interpolation='nearest', cmap=cmap)\n    plt.title(title)\n    plt.colorbar()\n    tick_marks = np.arange(len(classes))\n    plt.xticks(tick_marks, classes, rotation=45)\n    plt.yticks(tick_marks, classes)\n\n    fmt = '.2f' if normalize else 'd'\n    thresh = cm.max() / 2.\n    for i, j in itertools.product(range(cm.shape[0]), range(cm.shape[1])):\n        plt.text(j, i, format(cm[i, j], fmt),\n                 horizontalalignment=\"center\",\n                 color=\"white\" if cm[i, j] > thresh else \"black\")\n\n    plt.ylabel('True label')\n    plt.xlabel('Predicted label')\n    plt.tight_layout()\n    \ncnf_matrix = confusion_matrix(y_true, y_pred,labels=['bird','cat','dog','hen', 'pig'],)\nnp.set_printoptions(precision=2)\n\n# Plot non-normalized confusion matrix\nplt.figure()\nplot_confusion_matrix(cnf_matrix, classes=['bird','cat','dog','hen', 'pig'],\n                      title='Confusion matrix C, without normalization')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"### Example using our competition's dataset"},{"metadata":{},"cell_type":"markdown","source":"The above matrix is a **multi class confusion matrix**. As an example, for the 2nd row, we predicted 4 cats to be birds, and 1 cat to be predicted as hen. More compactly, it can be represented as the matrix $C_{2,1}$ = 4."},{"metadata":{},"cell_type":"markdown","source":"We can easily reconcile our example above to relate back to our competition: **After the biopsy is assigned a Gleason score, it is converted into an ISUP grade on a 1-5 scale**. For now, I will include the label **0** because it is the background/non-tissue. \n\nWhat I do next is to take the ground truth and call it `y_true` which is a series. I then generate a dummy `y_pred` by using `np.random.choice` and randomly generate numbers from 0 to 5."},{"metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:43:40.584317Z","start_time":"2020-04-22T04:43:40.566357Z"},"trusted":true},"cell_type":"code","source":"train = pd.read_csv(\"/kaggle/input/prostate-cancer-grade-assessment/train.csv\")\ny_true = train.isup_grade\ny_true\n\ny_pred = np.random.choice(6, 10616, replace=True)\ny_pred = pd.Series(y_pred)\ny_pred","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"The following confusion matrix, is what we mean by the \"N by N\" (6 by 6) **histogram matrix**. "},{"metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:43:43.243947Z","start_time":"2020-04-22T04:43:42.998584Z"},"trusted":true},"cell_type":"code","source":"cm = confusion_matrix(y_true, y_pred,labels=[0,1,2,3,4,5],)\nplot_confusion_matrix(cm, classes=[0,1,2,3,4,5],\n                      title='Confusion matrix C, without normalization')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"***So far we have settled the first portion, construction the histogram matrix.***"},{"metadata":{},"cell_type":"markdown","source":"## **Step 2: Create the Weighted Matrix w** <a id =\"weighted\"></a>"},{"metadata":{},"cell_type":"markdown","source":"An N-by-N matrix of weights, w, is calculated based on the difference between actual and predicted rating scores. The formula is as follows:  $$w_{i,j} = \\dfrac{(i-j)^2}{(N-1)^2}$$\n\n\nUnder Step-2, each element is weighted. Predictions that are further away from actuals are marked harshly than predictions that are closer to actuals. However, I suddenly realized that these animals above have no obvious hierarchy in them. Therefore, it is not correct to say that predicting a cat as a bird is any better/worse off than predicting a cat as hen.\n\n**Consequently, this is a realization. Because now I start to understand why certain competitions use metrics like quadratic weighted kappa!** Consider the same example as animals just now, but instead of animals, we change to **isup_grade**). So there is an inherent order within the **isup_grade** 0 - 5 such that 0 and 1 is closer than 0 and 2, 1 and 2 is closer to 1 and 3 etc. $$0 > 1 > 2 > 3 > 4 > 5$$\n\n<br>\n\nFor example, let's use a simplified example:\n\n    y_true = [2,2,2,1,2,3,4,5,0,1]\n    y_pred = [1,2,4,1,2,3,4,5,0,1]\n    \nAs a result, our purpose of the weighted matrix is to allocate a higher penalty score if our prediction is further away from the actual value. That is, if our **isup_grade** is 2 but we predicted it as 1 (see the example), then based on our formula above, we have i = 2 and j = 1 (entry $C_{2,1}$), the penalty is $$\\dfrac{(2-1)^2}{(5-1)^2}  = 0.0625$$ but if our **isup_grade** is 2 and we predicted it as 4 (see the example), then the penalty involved is higher $$\\dfrac{(2-4)^2}{(5-1)^2} = 0.25$$\n\nIndeed, the weight matrix helped us assign a heavier penalty to predicting 2 as 4 than 2 as 1.\n\n***Lastly, we also observe that the formula will give us 0 for the diagonals of the weighted matrix. This means penalty is 0 whenever we correctly predict something.***"},{"metadata":{},"cell_type":"markdown","source":"To calculate weighted matrix in python code, here is the code with reference to [Aman Arora](https://www.kaggle.com/aroraaman/quadratic-kappa-metric-explained-in-5-simple-steps)."},{"metadata":{"ExecuteTime":{"end_time":"2020-04-22T05:08:16.156024Z","start_time":"2020-04-22T05:08:16.152037Z"},"trusted":true},"cell_type":"code","source":"# We construct the weighted matrix starting from a zero matrix, it is like constructing a \n# list, we usually start from an empty list and add things inside using loops.\n\ndef weighted_matrix(N):\n    weighted = np.zeros((N,N)) \n    for i in range(len(weighted)):\n        for j in range(len(weighted)):\n            weighted[i][j] = float(((i-j)**2)/(N-1)**2) \n    return weighted\n        \nprint(weighted_matrix(5))","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"***Remember, the further away from the diagonal you get, the worse off is your prediction and it will be penalized harder by a bigger weight.***  As you can easily infer from the weighted matrix above, we use the first row as an example in case one did not understand. Basically, the weighted matrix's first row's first element is 0, because it means we predicted correctly and no penalty is meted out; but as we move further to the left, you can see that the punishment gets harsher and harsher: $$0 < 0.06 < 0.25 < 0.56 < 1$$"},{"metadata":{},"cell_type":"markdown","source":"## **Step 3: Create the Expected Matrix** <a id =\"expected\"></a>"},{"metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.669296Z","start_time":"2020-04-21T17:19:43.664309Z"},"trusted":true},"cell_type":"code","source":"## dummy example \nactual = pd.Series([2,2,2,3,4,5,5,5,5,5]) \npred   = pd.Series([2,2,2,3,2,1,1,1,1,3]) \nC = confusion_matrix(actual, pred)\n\nN=5\nact_hist=np.zeros([N])\nfor item in actual: \n    act_hist[item - 1]+=1\n    \npred_hist=np.zeros([N])\nfor item in pred: \n    pred_hist[item - 1]+=1\n    \n\nprint(f'Actuals value counts:{act_hist}, \\nPrediction value counts:{pred_hist}')","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"This part is the most difficult to understand, especially for someone who has little statistic backgrounds (mind you when I was majoring in applied math back then, I only took **one** statistic module in my whole tenure.) Googling the idea of expected matrix is not apparently clear to me, but luckily someone pointed to me to read expected frequency of Chi-Square test and then and, there I start to slowly understand.\n\nFor the purpose of better understanding, we call our actual values to be rater A and our prediction model to be rater B. We want to quantify the agreement between rater A and rater B. Here are some terminologies to get hold of first.\n\n- There are a total number of $k = 5$ classes in this example;\n- There are a total number of $n = 10$ observations in this example;\n- Define Y to be the random variable that rater A has chosen (aka our actual classes 1,2,3,4,5);\n- and $\\widehat{Y}$ be the random variable that rater $B$ has chosen (aka our predicted classes 1,2,3,4,5 by rater B)\n- $r_i$ be the i-th entry of the column vector for actual value counts shown above, $c_i$ be the i-th entry of the column vector for prediction value counts shown above.\n<br>\n\nThen the probability of rater A choosing class 2 and rater B choosing class 2 for example, is given by $$P(Y = 2 \\text{ and } \\widehat{Y} = 2)  = P(Y = 2) \\cdot P(\\widehat{Y} = 2) = 30\\% \\times 40\\%  = 12\\%$$\n\nThis is under the assumption that both raters are **independent of each other**. Note that $P(Y = 2) = 30\\%$ because as we see from the actual value counts of rater $A$, there are a total of 3 class 2's and therefore the probability of Y being 2 is just the proportion. Similarly, we calculate $\\widehat{Y}$ the same way.\n\nIn general, the formula of rater A choosing class i and rater B choosing (predicting) class j is given as follows: $$P(Y = i \\text{ and }  \\widehat{Y} = j) =  P(Y = i) \\times P (\\widehat{Y} = j)$$\n\n<br>\n\n\nNow the real question comes: On average, if you have $n$ number of points to predict, how many times (what is the frequency) would you **expect** to see rater A choose class i and rater B choose class j.\n\nTo reiterate, recall that the probability of the actual class being 1 **and (super important word here, it means a joint distribution)** and the predicted class to be 1 as well is $P(Y = 1 \\text{ and }  \\widehat{Y} = 1)$, similarly, the probability of the actual class being 1 and the predicted class to be 2 is $P(Y = 1 \\text{ and }  \\widehat{Y} = 2)$. Generalizing, the probability of the actual class being $i$ and the predicted class being $j$ is the joint probability $$P(Y = i \\text{ and }  \\widehat{Y} = j)$$\n\nSo the intuition lies here: based on the joint probability above, what is the expected number of times (frequency) that $Y = i$ and $\\widehat{Y} = j$ happened (this means Y is i but rater B predict $\\widehat{Y}$ as j) out of 10 times? Easy, just use $n \\times P(Y = i \\text{ and }  \\widehat{Y} = j)$. So for rater B, our prediction model, **by just using theoretical probability**, should have $n \\times P(Y = i \\text{ and }  \\widehat{Y} = j)$ for each $i,j$. But in reality, this may not be the case. Reconcile this idea with the classic coin toss example:\n\n- **Example on coin toss:** Expected frequency is defined as the number of times that we predict an event will occur based on a calculation using theoretical probabilities. You know how a coin has two sides, heads or tails? That means that the probability of the coin landing on any one side is 50% (or 1/2, because it can land on one side out of two possible sides). If you flip a coin 1,000 times, how many times would you expect to get heads? About half the time, or 500 out of the 1,000 flips. To calculate the expected frequency, all we need to do is multiply the total number of tosses (1,000) by the probability of getting a heads (1/2), and we get 1,000 * 1/2 = 500 heads. if event A has prob of p happening, and sample size of n, then the average or expected number of times that A happens is $np$. The expected frequency of heads is 500 out of 1,000 total tosses. The expected frequency is based on our knowledge of probability - we haven't actually done any coin tossing. However, if we had enough time on our hands, we could actually flip a coin 1,000 times to experimentally test our prediction. If we did this, we would be calculating the experimental frequency. The experimental frequency is defined as the number of times that we observe an event to occur when we actually perform an experiment, test, or trial in real life. If you flipped a coin 1,000 times and it landed on heads 479 times, then the experimental frequency of heads is 479.\n\n    But wait - didn't we predict the coin would land on heads 500 times? Why did it actually land on heads only 479 times? Well, the expected frequency was just that - what we expected to happen, based on our knowledge of probability. There are no guarantees when it comes to probability, what we expect to happen might differ from what actually happens. \n\n<br>\n\nSince we know $$E_{2,2} = 10 \\times P(Y = 2 \\text{ and } \\widehat{Y} = 2)  = 10 \\times P(Y = 2) \\cdot P(\\widehat{Y} = 2) = 10 \\times 30\\% \\times 40\\%  = 1.2$$\n\nThis $E_{2,2}$ means that we only expect the rater A to choose 2 and rater B to choose 2 at the same time only 1.2 times out of 10 times. In other words, our predicted model classify 2 as 2 1.2 times. However, from our observations in our matrix C (confusion/histogram matrix), rater A choosing 2 and rater B choosing 2 have a frequency of 3! In other words, our predicted model classified 2 as 2 three times! So we kinda exceeded expectation for this particular configuration.\n\n$$C = \\begin{bmatrix}\n0 & 0 & 0 & 0 & 0\\\\\n0 & 3 & 0 & 0 & 0 \\\\\n0 & 0 & 1 & 0 & 0\\\\\n0 & 1 & 0 & 0 & 0\\\\\n4 & 0 & 1 & 0 & 0\\\\\n\\end{bmatrix}$$\n\nLet me give you one more example, $$E_{5,2} = 10 \\times P(Y = 5 \\text{ and } \\widehat{Y} = 2) = 10 \\times P(Y = 5) \\cdot P(\\widehat{Y} = 2) = 10 \\times 50\\% \\times 40\\%  = 2$$\n\nThis means that we only expect the rater A to choose 5 and rater B to choose 2 at the same time only 2 times out of 10 times! But in reality, our observations say that our rater B classified 5 as 2 zero times! \n\n\nWe calculate $E_{i,j}$ given by the formula: $$E_{i,j} = n \\times P(Y = i \\text{ and }  \\widehat{Y} = j) = n \\times P(Y = i) \\times P (\\widehat{Y} = j) = n \\times \\dfrac{r_i}{n} \\times \\dfrac{c_j}{n}$$\n\n$$E = \\begin{bmatrix}\n0 & 0 & 0 & 0 & 0\\\\\n1.2 & 1.2 & 0.6 & 0 & 0 \\\\\n0.4 & 0.4 & 0.2 & 0 & 0\\\\\n0.4 & 0.4 & 0.2 & 0 & 0\\\\\n2 & 2 & 1 & 0 & 0\\\\\n\\end{bmatrix}$$\n\nNote $r_i \\times c_j$ is the $(i,j)$ entry of the outer product between the actual histogram vector of outcomes (actual value counts) and the predicted histogram vector (prediction value counts)."},{"metadata":{},"cell_type":"markdown","source":"What does the random chance mean here? Simple, it just means that the probability for rater A (actual) to be class i AND for rater B (predicted) to be class j is $p\\%$ (say 10%). Therefore, if we have made 100 predictions, random chance aka the theoratical probability tells us you should only have $100 \\times 10\\% = 10$ predictions to be of this configuration (rater A class i AND rater B class j)."},{"metadata":{},"cell_type":"markdown","source":"#### Writing out the expected matrix in python\n\nSo to get the expected matrix, E, is calculated assuming that there is no correlation between values.  This is calculated as the outer product between the actual histogram vector of outcomes and the predicted histogram vector, normalized such that E and C have the same sum. Note carefully below that we do not need to normalize both, we just need to normalize E to the same sum as C."},{"metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.752074Z","start_time":"2020-04-21T17:19:43.74609Z"},"trusted":true},"cell_type":"code","source":"E = np.outer(act_hist, pred_hist)/10\n\n\nE\nC","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## **Step 4: Final Step: Weighted Kappa formula and Its python codes** <a id=\"qwk\"></a>\n \nFrom these three matrices: E, C and weighted, the quadratic weighted kappa is calculated as: \n\n$$\\kappa = 1 - \\dfrac{\\sum_{i,j}\\text{weighted}_{i,j}C_{i,j}}{\\sum_{i,j}\\text{weighted}_{i,j}E_{i,j}}$$\n\nNote that a higher value generally means your prediction model is way better than a random model, but there is no consensus on which value is really good or bad."},{"metadata":{},"cell_type":"markdown","source":"$$\\text{Weighted} = \\begin{bmatrix}\n0 & 0.0625 & 0.25 & 0.5625 & 1\\\\\n0.0625 & 0 & 0.0625 & 0.25 & 0.5625 \\\\\n0.25 & 0.0625 & 0 & 0.0625 & 0.25\\\\\n0.5625 & 0.25 & 0.0625 & 0 & 0.0625\\\\\n1 & 0.5625 & 0.25 & 0.0625 & 0\\\\\n\\end{bmatrix}$$"},{"metadata":{},"cell_type":"markdown","source":"The notation $\\sum_{i,j}\\text{W}_{i,j}C_{i,j}$ is just $$\\sum_{i=1}^{k}\\sum_{j=1}^{k} W_{i,j} C_{i,j} = (W_{1,1}C_{1,1} + W_{1,2}C_{1,2} + ...+ W_{1,k}C_{1,k}) + (W_{2,1}C_{2,1}+...+W_{2,k}C_{2,k}) +...+(W_{k,1}C_{k,1}+...+W_{k,k}C_{k,k})$$\n\nTo put our understanding into perspective, consider just one entry $W_{5,1}C_{5,1} = 1 \\times 4 = 4$. This roughly means that our predicted model classified class 5 as class 1 FOUR times (re: $C_{5,1} = 4$), and since class 5 is so far away from class 1, we need to **punish** this wrong prediction more than the others. And we did see that the corresponding weight $W_{5,1} = 1$ is the highest weight.\n\nConsequently, the numerator being $\\sum_{i,j}\\text{W}_{i,j}C_{i,j}$ calculates the total \"penalty cost\" for the rater A (our predicted model), and similarly, $\\sum_{i,j}\\text{W}_{i,j}E_{i,j}$ calculates the total \"penalty cost\" for the rater B (our \"expected\" model). Therefore, you can think both values (num and den) as a cost function, and the lesser the better. And kappa formula tells us, if our $\\sum_{i,j}\\text{W}_{i,j}C_{i,j}$ is significantly smaller than $\\sum_{i,j}\\text{W}_{i,j}E_{i,j}$, this will yield a very small value of $$\\dfrac{\\sum_{i,j}\\text{weighted}_{i,j}C_{i,j}}{\\sum_{i,j}\\text{weighted}_{i,j}E_{i,j}}$$ which will yield a very high kappa value - signifying a better model."},{"metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.798949Z","start_time":"2020-04-21T17:19:43.793962Z"},"trusted":true},"cell_type":"code","source":"# Method 1\n# apply the weights to the confusion matrix\nweighted = weighted_matrix(5)\nnum = np.sum(np.multiply(weighted, C))\n# apply the weights to the histograms\nden = np.sum(np.multiply(weighted, E))\n\nkappa = 1-np.divide(num,den)\nkappa","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.80593Z","start_time":"2020-04-21T17:19:43.799946Z"},"trusted":true},"cell_type":"code","source":"# Method 2\n\nnum=0\nden=0\nfor i in range(len(weighted)):\n    for j in range(len(weighted)):\n        num+=weighted[i][j]*C[i][j]\n        den+=weighted[i][j]*E[i][j]\n\nweighted_kappa = (1 - (num/den)); weighted_kappa","execution_count":null,"outputs":[]},{"metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.812911Z","start_time":"2020-04-21T17:19:43.806927Z"},"trusted":true},"cell_type":"code","source":"# Method 3: Just use sk learn library\n\ncohen_kappa_score(actual, pred, labels=None, weights= 'quadratic', sample_weight=None)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Just a side note that one can also use SK learn's cohen_kappa_score function calculate the quadratic weighted kappa in this competition, with weights set to quadratic. Also please do refer to [CPMP's discussion topic for fast QWK computation](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145105)"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.7.6"},"latex_envs":{"LaTeX_envs_menu_present":true,"autoclose":false,"autocomplete":true,"bibliofile":"biblio.bib","cite_by":"apalike","current_citInitial":1,"eqLabelWithNumbers":true,"eqNumInitial":1,"hotkeys":{"equation":"Ctrl-E","itemize":"Ctrl-I"},"labels_anchors":false,"latex_user_defs":false,"report_style_numbering":false,"user_envs_cfg":false},"toc":{"base_numbering":1,"nav_menu":{},"number_sections":false,"sideBar":true,"skip_h1_title":false,"title_cell":"Table of Contents","title_sidebar":"Contents","toc_cell":false,"toc_position":{"height":"calc(100% - 180px)","left":"10px","top":"150px","width":"303.837px"},"toc_section_display":true,"toc_window_display":true},"varInspector":{"cols":{"lenName":16,"lenType":16,"lenVar":40},"kernels_config":{"python":{"delete_cmd_postfix":"","delete_cmd_prefix":"del ","library":"var_list.py","varRefreshCmd":"print(var_dic_list())"},"r":{"delete_cmd_postfix":") ","delete_cmd_prefix":"rm(","library":"var_list.r","varRefreshCmd":"cat(var_dic_list()) "}},"types_to_exclude":["module","function","builtin_function_or_method","instance","_Feature"],"window_display":false}},"nbformat":4,"nbformat_minor":4}