{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction\n\n**This is an ongoing notebook, attempting to describe (almost) all Machine Learning Metrics in details, with examples to illustrate. I will refer to [scikit-learn's official website](https://scikit-learn.org/) for reference, and quote the explanation verbatim, if necessary. This is also because they do provide a very detailed user-guide, but I believe not everyone has the time to go through all their examples, so let's leave this task for me to summarize them.** \n\n**We will start from basic metrics like `Accuracy`, `confusion-matrix` etc. We will also cover flaws of some metrics (ahem `Accuracy`), along with the pros.**\n\n**A reminder: This is a work in progress, please do expect changes here and there as time goes by, and improvements on the code/text quality will improve.**\n\n---\n\n<div class=\"alert alert-block alert-danger\">\n<b>Danger:</b> The implementation below are quite raw, which means they are not modularized on purpose. Often I took complicated `numpy` operations for granted, and when probed further on my understanding of a certain algorithm, I cannot answer in details because many details are abstracted in the `numpy` calls.\n</div>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import cohen_kappa_score, make_scorer\nimport itertools\nimport sklearn\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport random\nfrom typing import Dict, List\n\nfrom IPython.core.interactiveshell import InteractiveShell\nInteractiveShell.ast_node_interactivity = \"all\"","metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:38:04.461975Z","start_time":"2020-04-22T04:38:04.456987Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:06.348457Z","iopub.execute_input":"2021-09-20T10:37:06.348913Z","iopub.status.idle":"2021-09-20T10:37:06.353623Z","shell.execute_reply.started":"2021-09-20T10:37:06.348863Z","shell.execute_reply":"2021-09-20T10:37:06.352977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Performance Metrics and Scoring\n\nIn almost all of Kaggle's competitions, the evaluation metric is given to us. However, in reality, when we are solving real world problems in our company, we have to come up with the performance metric ourselves. \n\nConsider a training set, consisting of 1000 patients, in which we want to correctly classify these patients into whether they have cancer (positive class = 1) or no cancer (negative class = 0) based on some independent variables. Since it is a training set, we have the ground truth data and knows that only $\\dfrac{50}{1000}$ of the patients have cancer. We can then come up with a trivial classifier that says:\"Classify any patients as no cancer (0)\" and this trivial classifier will actually yield you $\\dfrac{950}{1000} = 95\\%$ accuracy. Does it really mean our trivial classifier is good just because it has an accuracy rate of $95\\%$? We shall see.\n","metadata":{}},{"cell_type":"markdown","source":"# **Metrics for Binary Classification**\n\nWe will first start off with understanding metrics in the Binary Classification setting, more specifically, our classifier/hypothesis/model $h$, takes in an input $X$ that details the features of a tumour, and subsequently outputs first a calibrated soft label (probability), and we use a decision threshold to classify this soft label into a hard label which in this binary setting, is either $0$ (benign) or $1$ (malignant). The above example resonates more to competitions like image classification to detect malignant tumours from images/scans.\n\n---\n\nNote that we take the positive class as $1$, which is malignant, and negative class as $0$, benign.","metadata":{}},{"cell_type":"markdown","source":"## Confusion Matrix\n\n### Definition\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> In binary classification (can be extended to multi class), a table of confusion (sometimes also called a confusion matrix), is a table with two rows and two columns that reports the number of false positives, false negatives, true positives, and true negatives. This allows more detailed analysis than mere proportion of correct classifications `Accuracy metric`. Accuracy is not a reliable metric for the real performance of a classifier, because it will yield misleading results if the data set is imbalanced (that is, when the numbers of observations in different classes vary greatly).\n</div>\n\n---\n\nA confusion matrix is an $N \\times N$ matrix, where $N$ is the number of classes being predicted. Confusion Matrix gives us a matrix as output and describes the complete performance of the model. In a Binary setting, $N = 2$, so the confusion matrix is $2 \\times 2$.\n\nThe correct predictions falls on the diagonal line of the matrix.\n\n4 important terms in Confusion Matrix:\n\n- True Positives : The cases in which we predicted YES/POSITIVE and the actual output was also YES/POSITIVE.\n- True Negatives : The cases in which we predicted NO/NEGATIVE and the actual output was NO/NEGATIVE.\n- False Positives : The cases in which we predicted YES/POSITIVE and the actual output was NO/NEGATIVE.\n- False Negatives : The cases in which we predicted NO/NEGATIVE and the actual output was YES/POSITIVE.\n\n---\n\nIf not mentioned otherwise, we will be talking about univariate and single predictions.\n\n![confusion matrix](https://raw.githubusercontent.com/reigHns/reighns-MLAlgorithms/master/reighns-metrics/data/images/Basic-Confusion-matrix.webp)","metadata":{}},{"cell_type":"markdown","source":"### True Positive (TP)\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> The ground truth is positive label and the predicted value from the classifier is also positive label.\n</div>\n","metadata":{}},{"cell_type":"markdown","source":"### True Negative (TN)\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> The ground truth is negative label and the predicted value from the classifier is also negative label.\n</div>","metadata":{}},{"cell_type":"markdown","source":"### False Positive (FP)\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> The ground truth is negative label and the predicted value from the classifier is positive label.\n</div>\n\n<div class=\"alert alert-block alert-info\">\n<b>Info:</b> This is also called the Type 1 Error.</div>\n","metadata":{}},{"cell_type":"markdown","source":"### False Negative (FN)\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> The ground truth is positive label and the predicted value from the classifier is negative label.\n</div>\n\n<div class=\"alert alert-block alert-info\">\n<b>Info:</b> This is also called the Type 2 Error.</div>\n","metadata":{}},{"cell_type":"markdown","source":"### When to use Confusion Matrix?\n\nYou use it everywhere. The Confusion matrix in itself is not a performance measure, but the information that it carries is so valuable that almost all the other classification metrics will need to refer to the confusion matrix.","metadata":{}},{"cell_type":"markdown","source":"### Implementation of Confusion Matrix","metadata":{}},{"cell_type":"code","source":"y_true = np.array([1,1,0,1,0,0])\ny_pred = np.array([1,1,1,0,0,0])\n\ndef reighns_confusion_matrix(y_true, y_pred):\n    tp, tn, fp, fn = 0, 0, 0, 0\n    \n    for y_t, y_p in zip(y_true, y_pred):\n        # if actual and predicted both are positive class\n        if y_t == y_p == 1:\n            tp += 1\n        # if actual and predicted both are negative class\n        elif y_t == y_p == 0:\n            tn += 1\n        # if actual is negative and predicted is positive\n        elif y_t == 0 and y_p == 1:\n            fp += 1\n        # if actual is positive and predicted is negative\n        elif y_t == 1 and y_p == 0:\n            fn +=1\n    \n    return tp, fp, tn, fn\n\ntp, fp, tn, fn = reighns_confusion_matrix(y_true, y_pred)\n\nprint('Outcome values : \\n', tp, fn, fp, tn)","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:06.355111Z","iopub.execute_input":"2021-09-20T10:37:06.355444Z","iopub.status.idle":"2021-09-20T10:37:06.370286Z","shell.execute_reply.started":"2021-09-20T10:37:06.355408Z","shell.execute_reply":"2021-09-20T10:37:06.369419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can check against `sklearn`.","metadata":{}},{"cell_type":"code","source":"# confusion matrix in sklearn\nfrom sklearn.metrics import confusion_matrix\n\n# confusion matrix, set positive class to be positive 1\nmatrix = confusion_matrix(y_true, y_pred, labels=[1,0])\nprint('Confusion matrix : \\n',matrix)\n\n# outcome values order in sklearn\ntp, fn, fp, tn = confusion_matrix(y_true, y_pred,labels=[1,0]).reshape(-1)\nprint('Outcome values : \\n', tp, fn, fp, tn)","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:06.371697Z","iopub.execute_input":"2021-09-20T10:37:06.371981Z","iopub.status.idle":"2021-09-20T10:37:06.389273Z","shell.execute_reply.started":"2021-09-20T10:37:06.371934Z","shell.execute_reply":"2021-09-20T10:37:06.388418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Accuracy**","metadata":{}},{"cell_type":"markdown","source":"### Definition\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> Formally, if $\\hat{y}^{(i)}$ is the predicted value of the i-th sample and the ground truth is $y^{(i)}$, then accuracy can be defined as the fraction of predictions that our classifier/hypothesis/model predicted correctly, over the total number of samples in question.\n</div>\n\n$$\\text{accuracy}(\\hat{y}^{(i)}, y^{(i)}) = \\dfrac{1}{\\text{num_samples}}\\sum_{i=1}^{\\text{num_samples}}\\mathrm{1}(y^{(i)} = \\hat{y}^{(i)}) = \\dfrac{\\text{Number of correctly classified cases}}{\\text{Number for all cases}}$$\n\nwhere $\\mathrm{1}(x)$ is the [indicator function](https://en.wikipedia.org/wiki/Indicator_function).\n\n---\n\nIn the Binary setting, we can also write the formula as such: $$\\text{Accuracy} = \\dfrac{TP+TN}{TP+TN+FP+FN}$$\n","metadata":{}},{"cell_type":"markdown","source":"### When to use Accuracy as a metric\n\nClasses are well balanced: Accuracy is a valid choice of evaluation for classification problems which are well balanced and not skewed or no class imbalance. Typically, one should plot EDA and see the classes - if they are roughly equal, then `accuracy` can be used.\n\n\n### When NOT to use Accuracy as a metric\n\n<div class=\"alert alert-block alert-danger\">\n<b>Danger:</b> \n<li> Classes are *severely* imbalanced.\n<li> In general, `accuracy` is a very basic metric, and may not tell you any more information that fits your business needs. However, understanding this easiest metric first is crucial to facilitate the next few sections. You almost never ever use `accuracy` as a single metric, see 2.5 for intuition.\n</div>\n\n---\n\n<div class=\"alert alert-block alert-warning\">\n<b>Example:</b> Consider an imbalanced set, where the training data set has 100 patients (data points), and the ground truth is 90 patients are of class = 0, which means that these patients do not have cancer, whereas the remaining 10 patients are in class 1, where they do have cancer. This is an example of class imbalance where the ratio of class 1 to class 0 is $1:9$.\n</div>   \n    \nConsider **a baseline (almost trivial) classifier**:\n\n```python\ndef baselineModel(patient_data):\n        training...\n    return benign\n```\n        \n\nwhere we predict the patient's class as the most frequent class. Meaning, the most frequent class in this question is the class = 0, where patients do not have cancer, so we just assign this class to everyone in this set. By doing this, we will inevitably achieve a **in-sample** accuracy rate of $\\frac{90}{100} = 90\\%$. But unfortunately, this supposedly high accuracy value is completely useless, because this classifier did not label any of the cancer patients correctly.\n\nThe consequence can be serious, assuming the test set has the same distribution as our training set, where if we have a test set of 1000 patients, there are 900 negative and 100 positive. Our model just literally predict every one of them as benign, yielding a $90\\%$ **out-of-sample** accuracy.\n\nWhat did we conclude? Well, for one, our `accuracy` can be 90% high and looks good to the laymen, but it failed to predict the most important class of people - yes, misclassifying true cancer patients as healthy people is an **SERIOUS OFFENCE**. \n\nAlso, intuitively speaking, most models (without regularization), ***will not be robust to*** imbalanced dataset. This means that if you were given a training set that is highly imbalanced, then your \"logistic regression\" might give you an overly confident accuracy score on you validation set (if the validation set is also imbalanced).","metadata":{"tags":[]}},{"cell_type":"markdown","source":"### Implementation of Accuracy","metadata":{}},{"cell_type":"code","source":"y_true = np.array([1,1,0,1,0,0])\ny_pred = np.array([1,1,1,0,0,0])\n\naccuracy_count = 0 # numerator\nnum_samples = len(y_true) # denominator\n\nfor y_t, y_p in zip(y_true, y_pred):\n    if y_t == y_p:\n        accuracy_count += 1\n\nreighns_accuracy = accuracy_count / num_samples\n\nprint(f\"hn accuracy: {reighns_accuracy}\")\n\nfrom sklearn.metrics import accuracy_score\n\nprint(f\"sklearn accuracy: {accuracy_score(y_true, y_pred, normalize=True)}\")","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:06.391192Z","iopub.execute_input":"2021-09-20T10:37:06.391764Z","iopub.status.idle":"2021-09-20T10:37:06.401454Z","shell.execute_reply.started":"2021-09-20T10:37:06.391707Z","shell.execute_reply":"2021-09-20T10:37:06.400443Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Precision**","metadata":{}},{"cell_type":"markdown","source":"### Definition\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> Precision measures how many of the samples predicted as positive are actually positive.\n</div>\n\n$$\\text{Precision} = \\dfrac{\\text{TP}}{\\text{TP} + \\text{FP}}$$\n\n---\n\nInformally, precision answers the question **what proportion of positive predictions was actually correct**? In other words, out of all the positive predictions made by the model, how many of those positive predictions were actually positive when compared to the ground truth?\n\n---\n\nWhen I learned this back then, it is not immediately obvious what the denominator is doing. Dissecting the formula helps. From the formula, one can understand that TP is **good**, but FP is **bad**. The dynamics is that as TP and FP are inversely related. Assuming a fixed threshold, the denominator is fixed as follows:\n\n$$\\text{Predicted Number of Positives} = \\text{TP} + \\text{FP}$$    \n\nThus minimizing FP is equivalent to maximizing TP, and both will lead to an increase in precision.\n\n---\n\n<div class=\"alert alert-block alert-danger\">\n<b>Danger:</b> One common source of confusion is that people think that Precision is only applied to the positive class. This is not true, even in binary classification, if the positive class is malignant and negative class is benign, we can calculate the precision as it is, and the precision score represents the score hinged upon the aforementioned labelling. If instead we label the positive class to be benign, and the negative class malignant, then we will get a different precision score because our definition of positive class changed! The same is applied to the Recall section.\n</div>","metadata":{}},{"cell_type":"markdown","source":"### When to use Precision?\n\n- When your company needs you to ***restrict the number of False Positives***. As this metric penalizes **False Positives**.\n\n- The formula is also suggesting that we are ignoring about False Negatives rates here, so your company does not care about FN.\n\n- In other words, as we will see later, there is a trade off between precision and recall, and restricting the number of FP ***may*** give rise to the increase in FN. So ultimately, bear in mind that it is not simply a matter of ***restricting the number of False Positives*** but a matter of ***in your business setting, whether achieving less FP is more important than achieving a less FN***.\n\n- In general, precision - recall curve, as we will see later, gives a better overview of 'accuracy' than `accuracy score`.\n\n### When NOT to use Precision?\n\n<div class=\"alert alert-block alert-danger\">\n<b>Danger:</b> \n<li> Notice that if a precision score is 1 for a class C where you treat class C as the positive class, you will know that $TP = TP + FP = 1 \\implies FP = 0$, there are 0 False Positives, but this does not tell us **anything** about the False Negatives.\n<li> When you prioritize recall/sensitivity more than precision for your business needs.\n<li> You almost never ever use `precision` as a single metric, see 2.5 for intuition.\n</div>\n\n\n\n<div class=\"alert alert-block alert-warning\">\n<b>Example:</b> Consider a pharmaceutical company named *Preggie* in China that has developed a new testing kit to detect pregnancy. The company tested its kits on 10,000 ladies in the Wuhan city - these 10,000 ladies are our training samples. There are two outcomes/classes, \n<li> positive class = pregnant\n<li> negative class = not pregnant\n\nThis clinical trial involving 10,000 ladies involves a random sampling (and we assume the classes are quite balanced). The Data Scientists in Preggie needs to come up with a binary classification model that predicts the \"accuracy\" of their testing kit. Of course, the data scientists in Preggie are smart, they evaluate their classifier/model not by their `accuracy` but by their `precision` instead. \n\nThis is because they want to minimize their False Positives in their classifier. From the company's perspective, they want their testing kit to **correctly identify as much positive pregnancy cases as possible**. If their testing kit has a lot of False Positives, it will drive many of their customers to make the wrong decision - like quickly getting married and buying houses. The company care less about False Negatives because you will eventually find out that you are pregnant even if the initial testing kit shows otherwise.\n</div>   \n\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"### Probabilistic Interpretation\n\nOne can also interpret precision and recall not as ratios but as [estimations of probabilities](https://en.wikipedia.org/wiki/Precision_and_recall), and assuming positive class is 1 and negative class is 0, and $\\hat{Y}$ is your estimate of ground truth $Y$.\n\n$ precision = \\frac{T_p}{T_p+F_p} = Pr(Y=positive | X\\hspace{1mm}is\\hspace{1mm}predicted\\hspace{1mm}as\\hspace{1mm}positive) = P(Y=1 | \\hat{Y} = 1)$\n\n- Precision is the estimated probability that a random point selected from the samples are positive. This might be a tough pill to swallow as someone who was never good in statistics. But if you try to think a bit further, you can form an intuition as follows:\n\n> If your classifier $h$ is trained and the last layer is say, sigmoid, which in binary classification, calibrates the logits and turn them into probabilities. I understand it this as given a randomly chosen point $x \\in X_{train}$, what is probability of this point $x$ to be positive given that it is predicted as positive by the classifer?","metadata":{}},{"cell_type":"markdown","source":"### Implementation of Precision","metadata":{}},{"cell_type":"code","source":"y_true = np.array([1,1,0,1,0,0])\ny_pred = np.array([1,1,1,0,0,0])\n\ndef reighns_confusion_matrix(y_true, y_pred):\n    tp, tn, fp, fn = 0, 0, 0, 0\n    \n    for y_t, y_p in zip(y_true, y_pred):\n        # if actual and predicted both are positive class\n        if y_t == y_p == 1:\n            tp += 1\n        # if actual and predicted both are negative class\n        elif y_t == y_p == 0:\n            tn += 1\n        # if actual is negative and predicted is positive\n        elif y_t == 0 and y_p == 1:\n            fp += 1\n        # if actual is positive and predicted is negative\n        elif y_t == 1 and y_p == 0:\n            fn +=1\n    \n    return tp, fp, tn, fn\n\ntp, fp, tn, fn = reighns_confusion_matrix(y_true, y_pred)\nreighns_precision = tp / (tp + fp)\n\nprint(f\"hn precision: {reighns_precision}\")\n\nfrom sklearn.metrics import precision_score\nsklearn_precision = precision_score(y_true, y_pred, average='binary')\n\nprint(f\"sklearn precision: {sklearn_precision}\")","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:06.403347Z","iopub.execute_input":"2021-09-20T10:37:06.403734Z","iopub.status.idle":"2021-09-20T10:37:06.414376Z","shell.execute_reply.started":"2021-09-20T10:37:06.403695Z","shell.execute_reply":"2021-09-20T10:37:06.41376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Recall/Sensitivity/True Positive Rate","metadata":{}},{"cell_type":"markdown","source":"### Definition\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> Recall measures: out of all the actual positives (the real cancer patients), how many of them were identified correctly by the classifier?\n</div>\n\n$$\\text{Recall}= \\dfrac{\\text{TP}}{\\text{TP} + \\text{FN}}$$\n\n---\n\nFrom the formula, we see the denominator to be defined as TP + FN, which is unsurprising as this gives you the actual number of positives.\n\n---\n\nAnother way to construct the formula can be easily seen when you have your end goal in mind. If your end goal is to correctly classify all actual positives as actual positives, that is, if the sample's ground truth label is positive, you want your classifier to predict all samples with positive label with as much accuracy as possible. And since \n$$\\text{Actual Number of Positives = TP + FN}$$ \nit is therefore easy to see that minimizing FN is also maximizing TP.","metadata":{}},{"cell_type":"markdown","source":"### When to use Recall?\n\n- When your company needs you to ***restrict the number of False Negatives***.\n\n- Recall/Sensitivity therefore quantifies the avoidance of False Negatives.\n\n### When to NOT use Recall?\n\n<div class=\"alert alert-block alert-danger\">\n<b>Danger:</b> \n<li> Notice that if a recall score is 1 for a class C where you treat class C as the positive class, you will know that $TP = TP + FN = 1 \\implies FN = 0$, there are 0 False Negatives, but this does not tell us **anything** about the False Positives.\n<li> When you prioritize precision more than recall for your business needs.\n<li> You almost never ever use `recall` as a single metric, see 2.5 for intuition.\n</div>\n\n---\n\n<div class=\"alert alert-block alert-warning\">\n<b>Example:</b> For disease data modeling, anything that doesn't account for false-negatives is a crime. Recall is a better measure than precision. Consider a pharmaceutical company named *PredictC* in China that has developed a new testing kit to detect whether a breast tumor is positive (malignant) or negative (benign). The company tested its kits on 10,000 ladies in the Wuhan city - these 10,000 ladies are our training samples. There are two outcomes/classes, \n<li> positive class = malignant\n<li> negative class = benign\n\nThis clinical trial involving 10,000 ladies involves a random sampling (and we assume the classes are quite balanced). The Data Scientists in Preggie needs to come up with a binary classification model that predicts the \"accuracy\" of their testing kit. Of course, the data scientists in Preggie are smart, they evaluate their classifier/model not by their `accuracy` but by their `recall` instead. \n\nThis is because they want to minimize their False Negatives in their classifier. From the company's perspective, they want their testing kit to **reduce the number of false negatives as much as possible**.\n</div>   ","metadata":{}},{"cell_type":"markdown","source":"### Probabilistic Interpretation\n\nOne can also interpret precision and recall not as ratios but as [estimations of probabilities](https://en.wikipedia.org/wiki/Precision_and_recall):\n\nWhen I look at this, it sounds like a $conditional$ probability to me -- probability that a test instance will be classified as positive, given that it is indeed positive:-\n\n$ recall = \\frac{T_p}{T_p+F_n} = Pr(X\\hspace{1mm}is\\hspace{1mm}predicted\\hspace{1mm}as\\hspace{1mm}positive | X=positive) = P(\\hat{Y}=1 | Y = 1)$\n\n> For intuition, see the same section in Precision.","metadata":{}},{"cell_type":"markdown","source":"### Implementation of Recall","metadata":{}},{"cell_type":"code","source":"y_true = np.array([1,1,0,1,0,0])\ny_pred = np.array([1,1,1,0,0,0])\n\ndef reighns_confusion_matrix(y_true, y_pred):\n    tp, tn, fp, fn = 0, 0, 0, 0\n    \n    for y_t, y_p in zip(y_true, y_pred):\n        # if actual and predicted both are positive class\n        if y_t == y_p == 1:\n            tp += 1\n        # if actual and predicted both are negative class\n        elif y_t == y_p == 0:\n            tn += 1\n        # if actual is negative and predicted is positive\n        elif y_t == 0 and y_p == 1:\n            fp += 1\n        # if actual is positive and predicted is negative\n        elif y_t == 1 and y_p == 0:\n            fn +=1\n    \n    return tp, fp, tn, fn\n\ntp, fp, tn, fn = reighns_confusion_matrix(y_true, y_pred)\nreighns_recall = tp / (tp + fn)\n\nprint(f\"hn recall: {reighns_recall}\")\n\nfrom sklearn.metrics import recall_score\nsklearn_recall = recall_score(y_true, y_pred, average='binary')\n\nprint(f\"sklearn recall: {sklearn_recall}\")","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:06.415399Z","iopub.execute_input":"2021-09-20T10:37:06.415837Z","iopub.status.idle":"2021-09-20T10:37:06.432651Z","shell.execute_reply.started":"2021-09-20T10:37:06.415779Z","shell.execute_reply":"2021-09-20T10:37:06.431926Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Coincidentally, I chose a bad example in which the accuracy, precision and recall are the same. In practice, this is rarely the case.","metadata":{}},{"cell_type":"markdown","source":"## The Precision-Recall Tradeoff\n\nDoes this term reminisce with the Bias-Variance Tradeoff? \n\n---\n\nMore specifically, when we talk about precision and recall in the sections above, we are fixated at one decision threshold of our classifier. One should note that both metrics are parametrized by $t$, the decision threshold. We can tune our threshold to achieve a better precision or recall, but usually not both, hence the tradeoff.\n\n---\n\nWe can read more from [Google's Machine Learning Crash Course on Precison and Recall](https://developers.google.com/machine-learning/crash-course/classification/precision-and-recall).","metadata":{}},{"cell_type":"markdown","source":"## Specificity/True Negative Rate\n\n$$TNR = \\dfrac{TN}{TN + FP} = P(\\hat{Y} = 0| Y=0) = 1 - FPR$$","metadata":{}},{"cell_type":"markdown","source":"## False Positive Rate\n\n<div class=\"alert alert-block alert-success\">\n    <b>Definition:</b> Out of all the real negative classes (negative ground truth), how many were predicted wrongly (predicted as positive from the model).\n</div>\n\n\n$$FPR = \\dfrac{FP}{FP + TN}=1-TNR$$","metadata":{}},{"cell_type":"markdown","source":"## F1-Score","metadata":{}},{"cell_type":"markdown","source":"### Intuition\n\nMotivated by the examples above, where using single precision or recall do not tell us much about the whole story. We thus turn to a combination of the above metrics. Penalizes extreme values of precision and recall more than arithmetic mean.","metadata":{}},{"cell_type":"markdown","source":"### Definition\n\n<div class=\"alert alert-block alert-success\">\n    <b>Definition:</b> The F1 score is the <b>harmonic mean</b> between Precision and Recall, bounded between 0 and 1.\n</div>\n\n---\n\n$$F1 = \\dfrac{2(\\text{Precision} \\times \\text{Recall})}{\\text{Precision} + \\text{Recall}}$$\n\n<div class=\"alert alert-block alert-warning\">\n<b>Example:</b> \nFor example, imagine that the blood protein levels in diseased people and healthy people are normally distributed with means of 2 g/dL and 1 g/dL respectively. A medical test might measure the level of a certain protein in a blood sample and classify any number above a certain threshold as indicating disease. The experimenter can adjust the threshold (black vertical line in the figure), which will in turn change the false positive rate. Increasing the threshold would result in fewer false positives (and more false negatives), corresponding to a leftward movement on the curve. The actual shape of the curve is determined by how much overlap the two distributions have.\n</div>   \n","metadata":{}},{"cell_type":"markdown","source":"## Receiver operating characteristic (ROC)","metadata":{}},{"cell_type":"markdown","source":"### Intuition\n\nMathematically, ROC graphs are two dimensional graphs in which the x-axis is the False Positive Rate (FPR) and the y-axis, the True Positive Rate (TPR). The curve is parametrized by the parameter $\\vec{thr}$ which represents the threshold of the classifier. The graph also depicts the tradeoffs between TPR and FPR, much like the dilemma of the Bias-Variance tradeoff. Also note that in the ROC space, each point on the graph represents a threshold, and therefore each point can have its own confusion matrix as well.\n\nWe come up with an example from the [Melanoma Competition](https://www.kaggle.com/c/siim-isic-melanoma-classification) where we denote a malignant cell to be 1, and benign to be 0. If we treat the **malignance class as positive class**, and the model you trained on outputs a probability vector (using softmax here) $[0.7, 0.3]$ corresponding to class 0 and 1 respectively, the value of 0.3 translates to saying that the image is 30% positive that it is malignant; in other words, it is 70% sure that this image is a benign cell. If we choose the default threshold to be the traditional $\\vec{thr}=0.5$, then the classifier will label this image as a $0$. This is because the thresholds defines our **hard label** from the **soft label**, and thus anything above the threshold 0.5, will be classified as a positive class 1, and negative class otherwise.\n\nIf however, you lower your classification threshold, say from 0.5 to 0.2, then our image will become now become positive class, indicating the image's cell to be malignant. Intuitively, the consequence is that more images will be classified to become positive as lowering the threshold will allow the model to predict true more often. The consequence is that the TPR will go up, and so will the FPR.\n\n---\n\n<div class=\"alert alert-block alert-warning\">\n<b>Example:</b> There are 10 ground truth targets of y_true = [1,1,1,0,0,0,0,0,1,0,0] and your model predicts y_pred = [0.6,0.7,0.4,0.6,0.55,0.4,0.3,0.2,0.6,0.1] which if you apply argmax to y_pred, then it will become y_pred_argmax = [1,1,0,1,1,0,0,0,1,0]. The TPR here is given by $\\frac{2}{4}$ since there are 4 positive ground truth, and among the predicted labels, the model correctly classify 2 positives correctly. The FPR is given by $\\frac{3}{6}$ because we gave 3 people the false alarm, predicting them to have cancer whereas they don't. Now if you lower you threshold to 0.2, then you can see that the new predicted label array to be [1,1,1,1,1,1,1,0,1,0] where the new calibrated TPR is $\\frac{3}{4}$ and the FPR is $\\frac{5}{6}$. Therefore, without any proofs, just intuition, one should be convinced that if you lower the threshold, more patients will be classified as positive, consequently, the TPR and FPR both increase. Conversely, if you increase the threshold, then the TPR and FPR will both decrease. This may not hold true in a monotone manner, as wrongly described earlier, as it can jolly well be the TPR or FPR do not change, as can be seen in the diagram in the section Ranking.\n</div>   \n\n","metadata":{}},{"cell_type":"markdown","source":"### Definition of ROC Curve\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> The ROC Curve is a graph that plots the True Positive Rate on the y-axis and False Positive Rate on the x-axis, furthermore, this curve is parametrized by a threshold vector $\\vec{t}$.\n</div>\n\n---\n\n> From Wikipedia: In binary classification, the class prediction for each instance is often made based on a continuous random variable $X$, which is a \"score\" computed for the instance (e.g. the estimated probability in logistic regression). Given a threshold parameter $T$, the instance is classified as \"positive\" if $X > T$, and \"negative\" otherwise. $X$ follows a probability density $f_1(x)$ if the instance actually belongs to class \"positive\", and $f_0(x)$ if otherwise. Therefore, the true positive rate is given by $TPR(T) = \\int_{T}^{\\infty}f_1(x)dx$ and the false positive rate is given by $FPR(T) = \\int_{T}^{\\infty}f_0(x)dx$. The ROC curve plots parametrically TPR(T) with FPR(T) as the varying parameter.\n\n<div class=\"alert alert-block alert-warning\">\n<b>Example:</b> \nFor example, imagine that the blood protein levels in diseased people and healthy people are normally distributed with means of 2 g/dL and 1 g/dL respectively. A medical test might measure the level of a certain protein in a blood sample and classify any number above a certain threshold as indicating disease. The experimenter can adjust the threshold (black vertical line in the figure), which will in turn change the false positive rate. Increasing the threshold would result in fewer false positives (and more false negatives), corresponding to a leftward movement on the curve. The actual shape of the curve is determined by how much overlap the two distributions have.\n</div>   \n","metadata":{}},{"cell_type":"markdown","source":"### Definition of Area under ROC Curve\n\n<div class=\"alert alert-block alert-success\">\n<b>Definition:</b> The AUROC is thus the area under the ROC Curve.\n</div>\n\nMore formally, in the [probabilistic perspective of AUC](https://www.alexejgossmann.com/auc/), AUC is the probability of a randomly chosen positive case outranks a randomly chosen negative case based on the classifier.\n\n$$AUC=P(f(x+)>f(x−)|\\text{class}(x+)=1, \\text{class}(x−)=0)=\\frac{1}{PN}\\sum_{i=1}^{P}\\sum_{j=1}^{N}1(f(x+)−f(x−))$$\n\nwhere \n\n- $f(x)$: classifier\n- P : # of true positive item,\n- N : # of true negative item\n\nIn other words, it measures how well the probability ranks based on their true classes. Thus, it is a threshold-invariant and scale-invariant metrics and only the sequence matters in the predicted probabilities. Based on this property, models with higher AUC indicate better discrimination between the two classes. However, the probabilities output from models with higher AUC don’t always generate well-calibrated probabilities. More information can be found here: [Safe Handling Instructions for Probabilistic Classification](https://www.youtube.com/watch?v=RXMu96RJj_s).","metadata":{}},{"cell_type":"markdown","source":"### AUROC as a Ranking \n\nOne confusing aspect of ROC space is the ranking system. This can be seen in the notebook I created here. Remember, if you code it out yourself from scratch, then it will be more beneficial as you can understand where ranking come into play (without using sklearn). The algorithm starts from the point where threshold is \\infty or in sklearn it starts with some other number. A threshold of infinity will guarantee that the point starts at (0,0). Thus, our very first point MUST start from the origin in this algorithm. Then assuming we do not consider an infinity number of thresholds, as this is too computationally expensive, we consider say 10 threshold values that we want to test (a common number is the number in the dataset). We divide the 10 values into 0.9,0.8,0.7,...,0.1, (for example only). Then we start from 0.9, the highest threshold, and move down to the lowest, in order (ranking). As we have seen just now, as you lower the threshold, both your TPR and FPR go up. Therefore, if you don't want to get your hands dirty, then the intuition is that if you have ground truth [0,1,1,0] and pred_1 = [0.03,0.99,0.05,0.06] and pred_2 = [0.15,0.92,0.89,0.91] then if you then imagine that your thresholds are given by thres_1 = [infinity,0.99,0.06,0.05,0.03] and thres_2 = [infinity, 0.92,0.91,0.89,0.15] , then you can calculate that the TPR and FPR rate at each of the threshold for both predictions are actually the same, consequently, forming the same ROC curve. (Consider plotting it). The idea here is we do not care what your values of the predictions are, in fact, in neural networks, transforming logits through softmax may not be a well calibrated (refer to my calibrated probability notes) probability anyways. We do however, care about the ranking, as you can see our thresholds are sorted in descending order, noticed that we only need that many thresholds for the dataset because only the thresholds at the predictions matter. If you take a number between 0.06 and 0.99 for the first threshold set, you will notice that between this threshold, the TPR and FPR will always be the same. Therefore, we conclude, without proof, that if two arrays of prediction has the exact same relative order, then the AUC for both predictions will be the same, which means that AUC is invariant to the scale of the predictions, and in fact invariant to any sort of transformation, that preserves the order (i.e. a non-negative linear transformation); (you can have numbers greater than 1 and the AUC will be the same try [100,200,150,160]).  A corollary of this is we can’t treat outputs of an AUC-optimized model as the likelihood that it’s true. Some models may be poorly calibrated (eg: its output is always between 0.3 and 0.32) but still achieve a good AUC score because its relative ordering is correct. This is something to look out for when blending together predictions of different models.\n\n<img src='https://raw.githubusercontent.com/reigHns/reighns-MLAlgorithms/master/reighns-metrics/data/images/ranking_auc.png' width=\"390\" height=\"200\"/>\n","metadata":{}},{"cell_type":"markdown","source":"Furthermore, if you see my notebook example, you can predict wrongly, but still have an AUC of 1. One last thing is about the predictions ordering, there is no rule that your predictions must SORT IN DESCENDING ORDER, for example: This will give you an AUC score of 1, even though it may not seem to predict everything correctly. Because the below order gives rise to the best AUC, which is 1 in this case, and hence this will give you 1 as well. If you switch a few numbers inside y_pred you will notice it can still stay at 1. However, if you reverse the list order, then you will get an AUC of 0 (the opposite of the best).\n\n```python\ny_true = [1, 1, 1, 1, 1, 1, 0, 0, 0, 0]\ny_pred = [0.99999, 0.98, 0.97, 0.96, 0.95, 0.94, 0.68139, 0.50961, 0.48880, 0.44951]\nfull_score_example = sklearn.metrics.roc_auc_score(y_true, y_pred)  \nprint(full_score_example) -> 1\n```","metadata":{}},{"cell_type":"code","source":"y_true = [1, 0, 1, 1, 0]\ny_pred = [0.5, 0.25, 0.2, 0.3, 0.1]\ny_pred_same_rank = [100, 25, 20, 30, 10]\n\n\nfpr_rank_1, tpr_rank_1, threshold_rank_1 = sklearn.metrics.roc_curve(y_true, y_pred, drop_intermediate=True)\nfpr_rank_2, tpr_rank_2, threshold_rank_2 = sklearn.metrics.roc_curve(y_true, y_pred_same_rank, drop_intermediate=True)\nroc_rank_1 = sklearn.metrics.roc_auc_score(y_true=y_true, y_score=y_pred, average=\"macro\", sample_weight=None, max_fpr=None)  # 0.833\nroc_rank_2 = sklearn.metrics.roc_auc_score(y_true=y_true, y_score=y_pred_same_rank, average=\"macro\", sample_weight=None, max_fpr=None)  # 0.833","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:06.433901Z","iopub.execute_input":"2021-09-20T10:37:06.434283Z","iopub.status.idle":"2021-09-20T10:37:06.445281Z","shell.execute_reply.started":"2021-09-20T10:37:06.434246Z","shell.execute_reply":"2021-09-20T10:37:06.444479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nplt.figure(figsize =[10,9])\nplt.title('Receiver Operating Characteristic')\nplt.plot(fpr_rank_1, tpr_rank_1, 'b', label = 'AUC = %0.2f' % roc_rank_1)\nplt.plot(fpr_rank_2, tpr_rank_2, 'b', label = 'AUC = %0.2f' % roc_rank_2)\nplt.legend(loc = 'lower right')\nplt.plot([0, 1], [0, 1],'r--')\nplt.xlim([0, 1])\nplt.ylim([0, 1])\nplt.ylabel('True Positive Rate')\nplt.xlabel('False Positive Rate')\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:06.530754Z","iopub.execute_input":"2021-09-20T10:37:06.531209Z","iopub.status.idle":"2021-09-20T10:37:06.822847Z","shell.execute_reply.started":"2021-09-20T10:37:06.531165Z","shell.execute_reply":"2021-09-20T10:37:06.82187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### ROC as C-Statistic\n\nROC can be interpreted as [c-statistics](https://stats.stackexchange.com/questions/193138/roc-curve-drawbacks).\nROC AUC has the property that it coincides with the $c$ statistic. The $c$ statistic measures the probability that a positive example is ranked higher than a negative example. In this sense, the ROC AUC answers the question of how well the model discriminates between the two classes.\n\nA model with high discrimination is not necessarily well calibrated. Suppose a logistic regression model predicts probabilities of 0.52 for positives and 0.51 for negatives (imagine 10 ground truth where 6 is positive and 4 is negative, then the author meant the associated probabilities with each of these ground truth is 0.52 and 0.51 respectively, for the positive and negative classes). This model has an AUC of 1 (recall you need not predict everything correctly to get an AUC of 1) but the probabilities aren't helpful in the sense of identifying which purported positives are highest-risk. Because all of the positives are assigned the same posterior probability, they can't be differentiated.\n\nMoreover, a well-calibrated model will have its maximum ROC AUC fixed by the ratio of positives to negatives in the data. This means that a model which has some very desirable probabilities (i.e. its posterior probabilities match the true probability) has a cap on its performance, and therefore an **uncalibrated** model could \"dominate\" in terms of ROC AUC.\n\nROC AUC doesn't tell you anything about the costs of different kinds of errors. For example, if you're trying to detect fraud, a 10,000 dollar purchase of uncertain provenance represents a larger potential loss than a 10 dollar purchase. But ROC AUC would treat both events as if they have the same weight -- obviously any reasonable model should be able to distinguish between these two types of error.\n\nROC AUC also tends to be dominated by the \"high FPR\" points. **Depending on the application, these points may be the least relevant.** Consider the case where the model is used to refer high-risk transactions to experts who will conduct further vetting. There may only be enough humans to assess 50 transactions per unit time; since the most highly-ranked transactions occur on the \"left hand\" size of the ROC curve by definition, this is also the region with the lowest area. So by looking at the whole AUC, you're optimistically biasing your results upwards, i.e. ROC AUC is buoyed by the observations \"to the right\" of the actual set of observations which humans will vet. (Illustration is simple. Draw a vertical line at FPR<0.5 on any ROC curve. The area to left is higher for all such vertical lines.) To avoid this, some people use partial ROC AUC, which has its own host of problems, chief among them that software implementations tend to assume that you're interested in truncation *at some value of FPR.* But in the case that you care about the top $n$ transactions, this approach is obviously wrong because the top $n$ transactions will happen at different FPR values for different classifiers. Standardization of partial AUC (to preserve the property that AUC < 0.5 is worse than random, 1 is perfect, 0 is worthless) incurs further difficulties.\n\nThe ROC curve itself is of little interest. \"Dominating\" classifiers can be assessed by AUC. Stochastic equivalence can be assessed by tests of equivalence of ranks. Prof. Harrell's comment drives at a consistent theme of his work, which is that the real question diagnostics should answer is one of risk assessment and utility optimization. Examining ROC AUC tends to encourage selection of truncation points, which should be avoided because it only provides partial information to decision makers.\n\nAlternative measures of performance (e.g. log-likelihood) characterize the calibration of the model and proper scoring rules generally have the quality that they encourage honest forecasts.","metadata":{}},{"cell_type":"markdown","source":"### Pros and Cons of AUROC\n\nBefore we go ham on the [Drawbacks of AUROC](https://stats.stackexchange.com/questions/193138/roc-curve-drawbacks), we first try to think the following:\n\n<div class=\"alert alert-block alert-info\">\n<b>Food For Thought:</b> I think intuitively you can say that if your model needs to perform equally well on the positive class as the negative class (for example, for classifying images between cats and dogs, you would like the model to perform well on the cats as well as on the dogs. For this you would use the AUROC.)<p>\n\nOn the other hand, if you're not really interested in how the model performs on the negative class, but just want to make sure every positive prediction is correct (precision), and that you get as many of the positives predicted as positives as possible (recall), then you should choose PRAUC (more with it later). For example, for detecting cancer, you don't care how many of the negative predictions are correct, you want to make sure all the positive predictions are correct, and that you don't miss any. (In fact, in this case missing a cancer would be worse than a false positive so you'd want to put more weight towards recall.)\n</div>","metadata":{}},{"cell_type":"markdown","source":"#### Pros: When your classes are more balanced\n\n\n\nROC curves are insensitive to changes in class distribution. Quote unquote from \"The analysis of ROC Curves\", we see that the **if the proportion of positive to negative instances changes in a test set, the ROC curves will not change**. This is because the AUC is equals to the probability of ranking a random positive example over a random negative example, and by definition this happens after you have drawn a positive and a negative, which indicates that we do not need to know anything about the original distribution and the class proportions.\n\n---\n\nAUROC curve better reflects the total amount of False Positives independent of in which class they come up. We can see this by a simple math example:\n\nEssentially, AUROC is measuring the TPR vs FPR ratio: We can interpret it as such\n\n$$TPR:FPR = \\dfrac{TP}{TP + FN} : \\dfrac{FP}{FP+TN} = \\dfrac{TP}{TP+FN} \\times \\dfrac{FP + TN}{FP} = \\dfrac{TP}{|+|} \\times \\dfrac{|-|}{FP}=\\dfrac{|-|}{|+|} \\times \\dfrac{TP}{FP}$$\n\n\n\n---\n\n<!-- This portion here is important and can be extended into the Cons section as well. In summary, if you notice that AUROC curve is made up by TPR vs FPR, and since FPR = 1- TNR, we can deduce the following:\n\n$$FPR = 1 - P(\\hat{Y}=0|Y=0) \\text{ and } TPR = P(\\hat{Y}=1|Y=1)$$\n\nNow you may wonder, why did I express them in this format? This is because we should understand it probabilistically. Notice that both TPR and FPR are **probabilities conditioned on the true class label**. As a result, AUROC curves is the same no matter what the baseline probability is. To make it more concrete, consider the sample set $\\mathcal{D}$ to have 100 samples, in which 90 is **negative** and 10 is **positive**. Thus the baseline classifier ZeroR will always predict majority, in this case, it will always predict negative.\n\n$$P(Y=1) = 10\\% \\text{ and } P(Y=0) = 90\\%$$ -->","metadata":{}},{"cell_type":"markdown","source":"#### Pros: Scale Invariant\n\nAUC measures how well predictions are ranked, rather than their absolute values. This can be a con as highlighted in Cons: uncalibrated.","metadata":{}},{"cell_type":"markdown","source":"#### Pros: Classification Threshold Invariant\n\nAUC measures the quality of the model's predictions irrespective of what classification threshold is chosen. What this means is if you compare an example to accuracy, how do you compute it? You say that if threshold is more than $t$, then you proceed to calculate the accuracy score - and different threshold gives different accuracies. But in ROC, the nuance is that our final metric is area under the ROC curve, over various (all possible) thresholds $t$, so as you see, we do not depend on the threshold to calculate the final score! \n\nThis can also be a con when you want to specifically minimize one metric like False Negatives or False Positives. For example, in cancer detection where malignant is the positive class, you will likely want to minimize False Negatives, **even if it results in a huge increase in False Positives**, then ROC may not be best suited. So if you only have the ROC curve for analysis, then you can choose your threshold according to the curve, in this case we choose the point which maximizes TPR as maximizing TPR is equivalent to minimizng FN.","metadata":{}},{"cell_type":"markdown","source":"#### Cons: Imbalanced\n\n<div class=\"alert alert-block alert-danger\">\n<b>Imbalanced Data:</b> We examine the case in which the dataset is imbalanced and further assume that the positive class is the minority, (note if you assume positive class is majority, then ROC may perform very well here, so the assumption is that the minority is of the positive class). We assume further that the negative class is 90% and positive class is 10%. Intuitively, in an imbalanced dataset, the model **usually does not have trouble predicting the majority class**, and this suggests that they will often get the negatives correct in this case, leading to a high TN. By looking at the FPR, we notice that $\\frac{FP}{FP + TN}$ suggests that FP will be low and TN will be high simply because of the aforementioned idea that the model will likely get the TN correct, and if TN is high, then the FP is low. Consequently, FPR is high. <p>\n\nThe following [on why AUC can be misleading](https://stats.stackexchange.com/questions/360017/when-is-an-auc-score-misleadingly-high/360040#360040)\nOne possible reason you can get high AUROC with what some might consider a mediocre prediction is if you have imbalanced data (in favor of the \"zero\" prediction), high recall, and low precision. That is, you're predicting most of the ones at the higher end of your prediction probabilities, but most of the outcomes at the higher end of your prediction probabilities are still zero. This is because the ROC score still gets most of its \"lift\" at the early part of the plot, i.e., for only a small fraction of the zero-predictions. \n\nFor example, if 5% of the test set are \"ones\" and all of the ones appear in the top 10% of your predictions, then your AUC will be at least 18/19 because, after 18/19 of the zeroes are predicted, already 100% of the ones were predicted. Even if the top 5% are all zeroes.\n\nWhether this is a \"bad\" prediction depends on your priorities. If you think that false negatives are terrible and false positives are tolerable, then this prediction is okay. But if it's the opposite, then this prediction is pretty bad.\n</div>","metadata":{}},{"cell_type":"code","source":"import sklearn\nimport numpy as np\nimport matplotlib.pyplot as plt\n\nyTest = [0,0,1,1,0,1,0,0,1,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0]\nyPredicted = np.linspace(0.9, 0.1, num=len(yTest))\nprint(yPredicted)\n\nimbalanced_roc = sklearn.metrics.roc_auc_score(yTest, yPredicted) # ~0.89\nprint(f\"roc score: {imbalanced_roc}\")\n\nfpr_imbalanced, tpr_imbalanced, threshold_imbalanced = sklearn.metrics.roc_curve(yTest, yPredicted, drop_intermediate=False)\n\nprint(fpr_imbalanced, tpr_imbalanced, threshold_imbalanced)\n\nplt.plot(fpr_imbalanced, tpr_imbalanced);","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:06.824861Z","iopub.execute_input":"2021-09-20T10:37:06.825379Z","iopub.status.idle":"2021-09-20T10:37:07.130696Z","shell.execute_reply.started":"2021-09-20T10:37:06.825322Z","shell.execute_reply":"2021-09-20T10:37:07.12955Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Cons: Uncalibrated\n\n<div class=\"alert alert-block alert-danger\">\n<b>Uncalibrated:</b> \nA model with high AUROC does not necessarily imply a well calibrated model. By this I mean a model with all extreme predictions for of say, it predicts all positive ground truths to be 0.52 and all negative ground truths to be 0.51. This model has an AUC of 1, but the probabilities aren't helpful in the sense of identifying which purported positives are highest-risk. Because all of the positives are assigned the same posterior probability, they can't be differentiated. Note that models like logistic regression are naturally well calibrated, but models like neural networks output logits, and hence we have to apply `sigmoid` or `softmax` to make it probabilities. <p>\n\nBelow is the \"phenomenon\" that AUC of 1 but the models look bad. Notice that even though we have a perfect AUROC score of 1, the model is not at all confident with the predictions in the sense that we cannot pin point any two positive labels and say that one of them is of higher probability than the other. If you are not convinced, the below code illustrates the point and the plot shows you.\n</div>","metadata":{}},{"cell_type":"code","source":"y_true = [1, 0, 1, 0, 1, 0, 1, 0, 1, 1]\n\ny_pred = [0.52, 0.51, 0.52, 0.51, 0.52, 0.51, 0.52, 0.51, 0.52, 0.52]\n\nuncalibrated_roc = sklearn.metrics.roc_auc_score(y_true, y_pred)\nprint(f\"roc score: {uncalibrated_roc}\")\n\nfpr_uncalibrated, tpr_uncalibrated, threshold_uncalibrated = sklearn.metrics.roc_curve(y_true, y_pred, drop_intermediate=True)\n\nprint(fpr_uncalibrated, tpr_uncalibrated, threshold_uncalibrated)\n\n# Another example\n# y_true = [1, 1, 1, 1, 1, 1, 0, 0, 0, 0]\n# y_pred = [0.99999, 0.98, 0.97, 0.96, 0.95, 0.94, 0.68139, 0.50961, 0.48880, 0.44951]\n# full_score_example = sklearn.metrics.roc_auc_score(y_true, y_pred)  \n# print(full_score_example) -> 1","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:07.132614Z","iopub.execute_input":"2021-09-20T10:37:07.133212Z","iopub.status.idle":"2021-09-20T10:37:07.149326Z","shell.execute_reply.started":"2021-09-20T10:37:07.132939Z","shell.execute_reply":"2021-09-20T10:37:07.147848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nsns.set()\nplt.figure(figsize=(15, 7))\n\nplt.scatter(fpr_uncalibrated, tpr_uncalibrated, color=\"#0F9D58\", s=100)\nplt.plot(fpr_uncalibrated, tpr_uncalibrated, color=\"#0F9D58\", label = f\"auc: {uncalibrated_roc}\")\nplt.title(\"ROC Curve\", fontsize=20)\nplt.legend(loc=4)\nplt.xlabel(\"False Positive Rate\", fontsize=16)\nplt.ylabel(\"True Positive Rate\", fontsize=16);","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:07.152271Z","iopub.execute_input":"2021-09-20T10:37:07.152754Z","iopub.status.idle":"2021-09-20T10:37:07.518232Z","shell.execute_reply.started":"2021-09-20T10:37:07.152655Z","shell.execute_reply":"2021-09-20T10:37:07.51691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Implementation of ROC and AUC\n\nThis implementation follows closely to [Implementation of ROC - roc-curve-and-auc-from-scratch-in-numpy-visualized - TDS](https://towardsdatascience.com/roc-curve-and-auc-from-scratch-in-numpy-visualized-2612bb9459ab).","metadata":{}},{"cell_type":"markdown","source":"#### Step 1: Problem Setup\n\n```python\ny_true_binary = np.asarray([0, 0, 1, 1, 0, 0, 1, 1])\ny_pred_binary = np.asarray([0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8])\n```\n\nWe have a binary classification problem with the targets and predictions shown above. We further note that the predictions are probabilities output from the **Sigmoid** layer in a logistic classifier.","metadata":{}},{"cell_type":"code","source":"y_true_binary = np.asarray([0, 0, 1, 1, 0, 0, 1, 1])\ny_pred_binary = np.asarray([0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8])","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:07.522115Z","iopub.execute_input":"2021-09-20T10:37:07.522511Z","iopub.status.idle":"2021-09-20T10:37:07.529443Z","shell.execute_reply.started":"2021-09-20T10:37:07.522436Z","shell.execute_reply":"2021-09-20T10:37:07.52785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 2: Define Threshold Range\n\nFor our classifier, our usual default threshold is as such:\n\n```python\nif y_pred_binary[i] > 0.5:\n    assign y_pred_binary as positive class (+)\nelse:\n    assign y_pred_binary as negative class (-)\n```\n\nThen it follows that different thresholds will result to different TPR and FPR. We can discretize our thresholds uniformly. Note that `scikit-learn` uses a different method to find the thresholds and are more optimized.\n\nFor starter, we will just set our threshold range from 0 to 1 with uniform interval of 0.1.","metadata":{}},{"cell_type":"code","source":"threshold_range = np.arange(0, 11, 1) / 10\nprint(threshold_range)","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:07.533823Z","iopub.execute_input":"2021-09-20T10:37:07.534673Z","iopub.status.idle":"2021-09-20T10:37:07.54432Z","shell.execute_reply.started":"2021-09-20T10:37:07.534241Z","shell.execute_reply":"2021-09-20T10:37:07.543588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 3: Classify prediction according to threshold\n\nThe next step we need to do is to classify our `y_pred_binary` from probabilities into hard labels, a 0 or 1 label. We create a dictionary `y_pred_thresholded` which has the threshold as key, and the value is the corresponding hard labels.\n\n```python\ny_pred_thresholded = {\n    0.0: [1, 1, 1, 1, 1, 1, 1, 1],\n    0.1: [1, 1, 1, 1, 1, 1, 1, 1],\n    0.2: [0, 1, 1, 1, 1, 1, 1, 1],\n    0.3: [0, 0, 1, 1, 1, 1, 1, 1],\n    0.4: [0, 0, 0, 1, 1, 1, 1, 1],\n    0.5: [0, 0, 0, 0, 1, 1, 1, 1],\n    0.6: [0, 0, 0, 0, 0, 1, 1, 1],\n    0.7: [0, 0, 0, 0, 0, 0, 1, 1],\n    0.8: [0, 0, 0, 0, 0, 0, 0, 1],\n    0.9: [0, 0, 0, 0, 0, 0, 0, 0],\n    1.0: [0, 0, 0, 0, 0, 0, 0, 0],\n}\n```","metadata":{}},{"cell_type":"code","source":"y_pred_thresholded: Dict = {}\n\nfor threshold in threshold_range:\n    \n    if threshold not in y_pred_thresholded:\n        y_pred_thresholded[threshold] = []\n\n    for y_p in y_pred_binary:\n        if y_p >= threshold:\n            y_pred_thresholded[threshold].append(1)\n        else:\n            y_pred_thresholded[threshold].append(0)\n","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:07.545379Z","iopub.execute_input":"2021-09-20T10:37:07.545729Z","iopub.status.idle":"2021-09-20T10:37:07.55594Z","shell.execute_reply.started":"2021-09-20T10:37:07.545693Z","shell.execute_reply":"2021-09-20T10:37:07.554944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 4: Calculate TPR and FPR \n\nNow we calculate the respective TPR and FPR for each thresholds's hard labels against the `y_true_binary`. We will make use of our `reighns_confusion_matrix` defined earlier to calculate.\n\n```python\ntpr_fpr = [\n    [1.0, 1.0],\n    [1.0, 1.0],\n    [1.0, 0.75],\n    [1.0, 0.5],\n    [0.75, 0.5],\n    [0.5, 0.5],\n    [0.5, 0.25],\n    [0.5, 0.0],\n    [0.25, 0.0],\n    [0.0, 0.0],\n    [0.0, 0.0],\n]\n```\n\nwhere the first element of the inner list is `tpr` and the second element is `fpr`.","metadata":{}},{"cell_type":"code","source":"tpr_fpr: Dict = {\"tpr\": [], \"fpr\": []}\n\nfor y_pred in y_pred_thresholded.values():\n    tp, fp, tn, fn = reighns_confusion_matrix(y_true_binary, y_pred)\n    tpr = tp / (tp + fn)\n    fpr = fp / (tn + fp)\n    tpr_fpr[\"tpr\"].append(tpr)\n    tpr_fpr[\"fpr\"].append(fpr)\n\nprint(tpr_fpr)","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:07.557242Z","iopub.execute_input":"2021-09-20T10:37:07.557493Z","iopub.status.idle":"2021-09-20T10:37:07.568433Z","shell.execute_reply.started":"2021-09-20T10:37:07.557438Z","shell.execute_reply":"2021-09-20T10:37:07.567411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 5: Plot the points as ROC Curve\n\nThe main idea of ROC Curve is to plot various pairs of [TPR, FPR] at different threshold on the graph, as shown below. Note that we **reversed** our tpr and fpr to be in line with Scikit-Learn. It does not affect the end result.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nsns.set()\nplt.figure(figsize=(15, 7))\n\ntpr_reighns = tpr_fpr[\"tpr\"][::-1]\nfpr_reighns = tpr_fpr[\"fpr\"][::-1]\n\nplt.scatter(fpr_reighns, tpr_reighns, color=\"#0F9D58\", s=100)\nplt.title(\"ROC Curve\", fontsize=20)\nplt.xlabel(\"False Positive Rate\", fontsize=16)\nplt.ylabel(\"True Positive Rate\", fontsize=16);","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:07.569694Z","iopub.execute_input":"2021-09-20T10:37:07.570044Z","iopub.status.idle":"2021-09-20T10:37:07.899003Z","shell.execute_reply.started":"2021-09-20T10:37:07.569887Z","shell.execute_reply":"2021-09-20T10:37:07.897922Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us compare to the `scikit-learn's` version!","metadata":{}},{"cell_type":"code","source":"\nfrom sklearn.metrics import roc_curve\nfpr_sklearn, tpr_sklearn, thresholds_sklearn = roc_curve(y_true_binary, y_pred_binary)\n\nplt.figure(figsize=(15, 7))\nplt.scatter(fpr_sklearn, tpr_sklearn, s=100, alpha=0.5, color=\"blue\", label=\"Scikit-learn\")\nplt.scatter(fpr_reighns, tpr_reighns, color=\"#0F9D58\", s=100, alpha=0.3, label=\"Our implementation\")\nplt.title(\"ROC Curve\", fontsize=20)\nplt.xlabel(\"False Positive Rate\", fontsize=16)\nplt.ylabel(\"True Positive Rate\", fontsize=16)\nplt.legend();","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:07.90041Z","iopub.execute_input":"2021-09-20T10:37:07.900752Z","iopub.status.idle":"2021-09-20T10:37:08.272177Z","shell.execute_reply.started":"2021-09-20T10:37:07.900699Z","shell.execute_reply":"2021-09-20T10:37:08.271182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Notice that in Scikit-Learn's version, they have 3 less points that us, this is discussed in details in the reference links I appended below. But just know that the end result is the same when we go to AUC!","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-danger\">\n<b>Starting and Ending Point:</b> \nNotice that the starting point and ending point of the ROC curve always start with (0, 0) and (1, 1). See the `y_pred_thresholded` we got earlier.\n</div>\n\n```python\ny_true_binary = np.asarray([0, 0, 1, 1, 0, 0, 1, 1])\n\ny_pred_thresholded = {\n    0.0: [1, 1, 1, 1, 1, 1, 1, 1],\n    0.1: [1, 1, 1, 1, 1, 1, 1, 1],\n    0.2: [0, 1, 1, 1, 1, 1, 1, 1],\n    0.3: [0, 0, 1, 1, 1, 1, 1, 1],\n    0.4: [0, 0, 0, 1, 1, 1, 1, 1],\n    0.5: [0, 0, 0, 0, 1, 1, 1, 1],\n    0.6: [0, 0, 0, 0, 0, 1, 1, 1],\n    0.7: [0, 0, 0, 0, 0, 0, 1, 1],\n    0.8: [0, 0, 0, 0, 0, 0, 0, 1],\n    0.9: [0, 0, 0, 0, 0, 0, 0, 0],\n    1.0: [0, 0, 0, 0, 0, 0, 0, 0],\n}\n```\n\nA bit of probing reveals that if you discretize your threshold from 0 to 1 inclusive, then it follows that at the threshold 1, everything is predicted as the negative class as shown, then by definition, TPR is 0 because the numerator of TPR is TP, and there is 0 TP because every single prediction made is of negative class, similarly, FPR is also 0 because the numerator of FPR is FP, and the model did not miss any negatives since it predicted every single one as negative.\n\n---\n\nThe same logic can be applied to when the threshold is 0, we instead have FPR and TPR to be both 1.","metadata":{}},{"cell_type":"markdown","source":"#### Step 6: Area under ROC Curve\n\nIn the [Implementation of ROC - roc-curve-and-auc-from-scratch-in-numpy-visualized - TDS](https://towardsdatascience.com/roc-curve-and-auc-from-scratch-in-numpy-visualized-2612bb9459ab) visualization, we understand that we can approximate AUROC score by \"integrating\" over the rectangles. In a way, this is calculating areas of rectangles under the curve, as follows.","metadata":{}},{"cell_type":"code","source":"rectangle_area = 0\nfor k in range(len(threshold_range)):\n    rectangle_area += (fpr_reighns[k+1] - fpr_reighns[k]) * tpr_reighns[k]\n    if k == (len(threshold_range) - 2): \n        break","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.273707Z","iopub.execute_input":"2021-09-20T10:37:08.274229Z","iopub.status.idle":"2021-09-20T10:37:08.280438Z","shell.execute_reply.started":"2021-09-20T10:37:08.274168Z","shell.execute_reply":"2021-09-20T10:37:08.279511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f\"reighns roc_auc_score: {rectangle_area}\")\n\nfrom sklearn.metrics import roc_auc_score\nroc_auc_sklearn = roc_auc_score(y_true_binary, y_pred_binary)\nprint(f\"sklearn roc_auc_score: {roc_auc_sklearn}\")","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.281917Z","iopub.execute_input":"2021-09-20T10:37:08.282425Z","iopub.status.idle":"2021-09-20T10:37:08.296249Z","shell.execute_reply.started":"2021-09-20T10:37:08.282368Z","shell.execute_reply":"2021-09-20T10:37:08.294961Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# alternatively, we can use np.trapz to calculate the area under the curve.\nroc_area = np.trapz(y=tpr_reighns, x=fpr_reighns)\nprint(roc_area)","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.297762Z","iopub.execute_input":"2021-09-20T10:37:08.298192Z","iopub.status.idle":"2021-09-20T10:37:08.305877Z","shell.execute_reply.started":"2021-09-20T10:37:08.298029Z","shell.execute_reply":"2021-09-20T10:37:08.304346Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Summary**\n\n1. AUC is a threshold-free metrics capable of measuring the overall performance of binary classifier.\n2. AUC should be used in binary classification. In multinomial classification, one-to-rest AUC would be an option using the average of each class.\n3. AUC is a good metric when the rank of output probabilities is of interest.\n4. Note that if you have 2 classes, then finding the AUROC of the positive class (class 1) is equivalent to 1 minus the AUROC of the negative class (class 0). This is not true when we deal with PR-curve.\n5. **Although AUC is powerful, it is not a cure-all. AUC is not suitable for heavily imbalanced class distribution and when the goal is to have well-calibrated probabilities.**\n6. Models with maximized AUC treat the weight between positive and negative class equally.\n7. AUROC would be the metric to use if the goal of the model is to perform equally well on both classes. Image classification between cats & dogs is a good example because the performance on cats is equally important on dogs. \n\n    AUPRC would be the metric to use if the focus of the model is to identify correctly as many **positive** samples as possible. Take spam detectors for example, the goal is to find all the possible spams. Regular emails are not of interest at all — they overshadow the number of positives.\n\n    There are no defined rules to select the suitable metrics. It really depends on the data and the application. It is important to think thoroughly about the purpose of the model before jumping into the modeling process.\n\n    One thing to note here is that the PR AUC serves as an alternative metric. If the model doesn’t work after the metric is changed, there are still other remedies to deal with imbalanced data, such as downsampling/upsampling. We’ll cover it later in future posts.","metadata":{}},{"cell_type":"markdown","source":"\n\n### SKLEARN Definition of Binary Classification ROC-AUC\n\n\n    sklearn.metrics.roc_auc_score(y_true, y_score, *, average='macro', sample_weight=None, max_fpr=None, multi_class='raise', labels=None)\n    \n    \n**y_score: array-like of shape (n_samples,) or (n_samples, n_classes)**\n\nTarget scores. In the binary and multilabel cases, these can be either probability estimates or non-thresholded decision values (as returned by decision_function on some classifiers). In the multiclass case, these must be probability estimates which sum to 1. The binary case expects a shape (n_samples,), and the scores must be the scores of the class with the greater label. The multiclass and multilabel cases expect a shape (n_samples, n_classes). In the multiclass case, the order of the class scores must correspond to the order of labels, if provided, or else to the numerical or lexicographical order of the labels in y_true.\n\nUnderstanding the binary case is important, it says that the binary case expects a list/array of shape `(n_samples,)`, a 1d-array, where the scores inside the 1d-array must be the scores of the **greater label**. In other words, if you have class 0 and 1, then the greater label is `np.argmax(0,1) = 1`. As a consequence, it is important that you should only pass the \"positive class\" which is the \"greater label\" here into the `y_score`. \n\nIn multiclass, there are two cases, either you provide a `labels` argument in, say `labels = [0,2,1]` or `labels = [0,1,2]`, or if you do not provide, then the `y_score` will necessarily be in the order of the numerical/alphabetical order of the labels in `y_true`. In other words, if `y_true` has 3 unique labels: 0, 1 and 2; then the `y_score` will be a **2d-array** in the form of `y_score = [[0.2, 0.3, 0.5],[...],[...]]` where `y_score[0] = [0.2,0.3,0.5]` must correspond to class 1, 2 and 3 respectively, unless otherwise stated in `labels`.\n\n\n### First Interpretation\n\nNow ROC curve is a TPR vs FPR graph, and the AUC is the area under the curve literally. To find the ROC-AUC, we need to plot many different pairs of points on the graph, and compute the area under it.\n\nAs we can see from the above naive and simple example, there are a total of 6 pairs of points to plot. Those are from `fpr`, `tpr` respectively --> Allow me to further explain with this example where 1 is the positive class:\n\n    y_true_1 = [0,0,1,1,0,0,1,1]\n    y_preds_1 = [0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]\n\n1. We need to initialize the thresholds with a large number usually - usually `roc_curve` is written so that ROC point corresponding to the highest threshold `(fpr[0], tpr[0])` is always `(0, 0)`. If this is not the case, a new threshold is created with an arbitrary value of `max(y_score)+1`. Therefore, in this case, we get 1.8 as the first threshold. This large number will ensure the `fpr, tpr` starts at (0,0).\n\n2. Next, when the threshold is $T=0.8$, then one can see that `y_preds_1` has 1 predictions 1, so `y_preds_1=[0,0,0,0,0,0,0,1]` and hence we can calculate the FPR and TPR: FPR will be 0 because no negative samples 0 are misclassified as 1 in our prediction. TPR will be 0.25 because by definition TPR=TP/TP+FN = 1/1+3=0.25 by definition. Therefore `(fpr, tpr) = (0,0.25)`\n\n3. $T=0.7 \\rightarrow$ `y_preds_1 = [0,0,0,0,0,0,1,1]`, same logic, FPR will be 0 cause no negative samples 0 are classified as 1 by our classifier! But TPR will be 0.5 because TPR = TP/TP+FN = 2/2+2 = 0.5. Therefore `(fpr, tpr) = (0,0.5)`\n\n4. We continue this way until we exhaust all thresholds given `[array([1.8, 0.8, 0.7, 0.5, 0.3, 0.1])]`. And we plot on the graph. \n\n5. How then, do we calculate the area under this curve? One can refer to the source code `auc` in `sklearn.metrics.auc` and see that they used **[Trapezoidal Rule](https://en.wikipedia.org/wiki/Trapezoidal_rule)** to solve it.\n\n6. So one have a rough idea, how the `ROC-AUC` area is computed, and one has to bear in mind that the area is calculated over all thresholds (apparently not the case as `sklearn` discretized the thresholds to reduce computing time, so you will not see the full range of thresholds here). \n\n\n","metadata":{}},{"cell_type":"markdown","source":"## **Precision-Recall Curve**","metadata":{}},{"cell_type":"markdown","source":"### Loss function and Decision Function\n\n[Loss function and decision Function link](https://stats.stackexchange.com/questions/104988/what-is-the-difference-between-a-loss-function-and-decision-function)\n\nA **decision function** is a function which takes a dataset as input and gives a decision as output. What the decision can be depends on the problem at hand. Examples include:\n\n - *Estimation problems:* the \"decision\" is the estimate.\n - *Hypothesis testing problems:* the decision is to reject or not reject the null hypothesis. Think of Linear Regression problems, they are mostly related to hypothesis testing.\n - *Classification problems:* the decision is to classify a new observation (or observations) into a category.\n - *Model selection problems:* the decision is to chose one of the candidate models.\n\nTypically, there are an infinite number of decision functions available for a problem. If we for instance are interested in estimating the height of Swedish males based on ten observations $\\mathbf{x}=(x_1,x_2,\\ldots,x_{10})$, we can use any of the following decision functions $d(\\mathbf{x})$:\n\n - The sample mean: $d(\\mathbf{x})=\\frac{1}{10}\\sum_{i=1}^{10}x_i$.\n - The median of the sample: $d(\\mathbf{x})=\\mbox{median}(\\mathbf{x})$\n - The geometric mean of the sample: $d(\\mathbf{x})=\\sqrt[10]{x_1\\cdots x_{10}}$\n - The function that always returns 1: $d(\\mathbf{x})=1$, regardless of the value of $\\mathbf{x}$. Silly, yes, but it is nevertheless a valid decision function.\n\nHow then can we determine which of these decision functions to use? One way is to use a **loss function**, which describes the loss (or cost) associated with all possible decisions. Different decision functions will tend to lead to different types of mistakes. The loss function tells us which type of mistakes we should be more concerned about. The best decision function is the function that yields the lowest **expected loss**. What is meant by expected loss depends on the setting (in particular, whether we are talking about [frequentist][1] or [Bayesian][2] statistics).\n\n**In summary:** \n\n - Decision functions are used to make decisions based on data.\n - Loss functions are used to determine which decision function to use.\n\n\n  [1]: http://en.wikipedia.org/wiki/Frequentist_inference\n  [2]: http://en.wikipedia.org/wiki/Bayesian_statistics\n  \n","metadata":{}},{"cell_type":"markdown","source":"### An extensive study on Precision-Recall Curve\n\nBefore I start, I will quote [Frank Harrell's first](https://www.fharrell.com/post/classification/) and [his second](http://www.fharrell.com/2017/03/damage-caused-by-classification.html) article. I will also use the links on stack exchange [here](https://stats.stackexchange.com/questions/312780/why-is-accuracy-not-the-best-measure-for-assessing-classification-models) and [here](https://stats.stackexchange.com/questions/368949/example-when-using-accuracy-as-an-outcome-measure-will-lead-to-a-wrong-conclusio) by [Stephan Kolassa](https://stats.stackexchange.com/users/1352/stephan-kolassa). They are really good and the entire intuition is from him, I will almost use his intuition verbatim and everything in this section will be credited to the links above; I just find it too difficult to phrase it in my own words because their answers are perfect.\n\nBasically, let me put it up front now: ***Accuracy, Sensitivity and Specificity are one-sided or conditional versions of classification accuracy. As such they are also discontinuous improper accuracy scores, and optimizing them will result in the wrong model.***  \n\n\n\nThe `confusion matrix` and the `classification report` provide a very detailed analysis of a particular set of predictions. However, the predictions themselves already threw away a lot of information that is contained in the model - to explain this statement further: We consider the example of a logistic regression classifier, used to predict whether a patient has cancer (1, positive class) or not (0, negative class). We defined Y as our response variable, outputting only 1 or 0, while $X$ is the set of predictors.\n\nIn the case of our cancer classification model, (which we assume to be a logistic regression classifier), we remember that the positive class (class = 1) is the patient has cancer, and the negative class (class = 0) is the patient does not have cancer. And to delve a little deeper, our default classification threshold is:\n    \n$\\begin{equation}\n Y=\\begin{cases}\n1, & \\text{if $P(Y=1 ~|~X) \\geq 0.5$}\\\\\n0, & \\text{if $P(Y=1~|~X) < 0.5$}\\\\\n\\end{cases} \\end{equation}$\n\nwhich means that whenever our logistic regression outputs a probability of the patient getting cancer is more than $0.5$, we classify the patient to be in the positive class (predict him/her to have cancer). When we use the `LogisticClassifier()` to fit and predict, we are actually predicting the probability $p(X)$, i.e.the probability of the patient having cancer given predictors X; Consequently, we need to further set a threshold, or to make a decision on whether to classify a patient as cancer or benign based on the probability we get from $$p(X) = \\dfrac{e^{\\beta_0}+\\beta_1X_1+...+\\beta_nX_n}{1+ e^{\\beta_0}+\\beta_1X_1+...+\\beta_nX_n}$$\n\nThe threshold is defaulted to 0.5 in `predict_proba`. As we discussed earlier, most classifiers provide a `decision_function` or a `predict_proba` method to assess degrees of certainty about predictions. Making predictions can be seen as thresholding the output of `decision_function` or `predict_proba` at a certain fixed point; in binary classification we use 0 for the decision function and 0.5 for predict_proba as default.\n\n\n<br>\n\n\nTo fully evaluate the effectiveness of a model, you must examine both precision and recall. Unfortunately, precision and recall are often in tension. That is, improving precision typically reduces recall and vice versa. Explore this notion by looking at the following figure, which shows 30 predictions made by an email classification model. Those to the right of the classification threshold are classified as \"spam\", while those to the left are classified as \"not spam.\"\n\nAlways remember, do not ever just use a single metric like `recall, precision` to gauge your classifier. This is because your classifier (say `SVM()` may somehow trivially classify everything as the positive class,  and then you will get 100% recall).","metadata":{}},{"cell_type":"markdown","source":"### When to use Precision-Recall\n\n[Precision-Recall curves should be used when there is a moderate to large class imbalance.](http://pages.cs.wisc.edu/~jdavis/davisgoadrichcamera2.pdf). In particular, when the positive class\n\nPrecision and recall, however, don't consider true negatives and thus won't be affected by the relative imbalance (which is precisely why they're used for imbalanced datasets). \n\n\n[Kaggle Forum](https://www.kaggle.com/general/7517#41179)\n\nAs goes for any metric, your metric depends entirely on what I you mean to do with the data.\n\nI think intuitively you can say that if your model needs to perform equally well on the positive class as the negative class (for example, for classifying images between cats and dogs, you would like the model to perform well on the cats as well as on the dogs. For this you would use the ROC AUC.\n\nOn the other hand, if you're not really interested in how the model performs on the negative class, but just want to make sure every positive prediction is correct (precision), and that you get as many of the positives predicted as positives as possible (recall), then you should choose PR AUC. For example, for detecting cancer, you don't care how many of the negative predictions are correct, you want to make sure all the positive predictions are correct, and that you don't miss any. (In fact, in this case missing a cancer would be worse then a false positive so you'd want to put more weight towards recall.)\n\nTrue negatives need to be meaningful for ROC to be a good choice of measure. In his example, if we've got 1,000 pictures of cats and dogs and our model determines whether the picture is a cat (target = 0) or a dog (target = 1), we probably care just as much about getting the cats right as the dogs, and so ROC is a good choice of metric.\n\nIf instead, we've got a collection of 1,000,000 pictures and we build a model to try to identify the 1,000 dog pictures mixed in it, correctly identifying \"not-dog\" pictures is not quite as useful. Instead, it makes more sense to measure how often a picture is a dog when our model says it's a dog (i.e., precision) and how many of the dogs in the picture set we found (i.e., recall).\n\n\n\n\n<div class=\"alert alert-block alert-info\">\n<b>Perspective:</b> In the cancer example above, your AUROC score might be very bad, simply because your False Positives might be high, as a result of minimizing False Negatives, but your AUPRC might be good, because you are maximizing precision!\n</div>","metadata":{}},{"cell_type":"markdown","source":"### When NOT to use Precision-Recall\n\n<div class=\"alert alert-block alert-danger\">\n<b>Majority Negative:</b> \nNotice that PR curve does not have TN in their equations, and this implies that PR curves are useful when there are minority positive samples and majority negative samples. But if it is the other way round, with minority negative samples, then PR curve will not tell you useful things.\n</div>","metadata":{}},{"cell_type":"markdown","source":"### Implementation of PR-Curve","metadata":{}},{"cell_type":"markdown","source":"## The Debate: AUROC vs AUPRC\n\nI just finished reading this discussion. They argue that PR AUC is better than ROC AUC on imbalanced dataset.\n\nFor example, we have 10 samples in test dataset. 9 samples are positive and 1 is negative. We have a terrible model which predicts everything positive. Thus, we will have a metric that TP = 9, FP = 1, TN = 0, FN = 0.\n\nThen, Precision = 0.9, Recall = 1.0. The precision and recall are both very high, but we have a poor classifier.\n\nOn the other hand, TPR = TP/(TP+FN) = 1.0, FPR = FP/(FP+TN) = 1.0. Because the FPR is very high, we can identify that this is not a good classifier.\n\nClearly, ROC is better than PR on imbalanced datasets. Can somebody explain why PR is better?\n\n---\n\nUsually when I do imbalanced models, even balanced models, I look at PR for ALL my classes.\n\nIn your example, yes, your positive class has P = 0.9 and R = 1.0. But what you should look at are ALL your classes. So for your negative class, your P = 0 and your R = 0. And you usually don't just look at PR scores individually. You want to look at F1-score (F1 macro or F1 micro, depending on your problem) that is a harmonic average of your PR scores for both class 1 and class 0. Your class 1 PR score is super good, but combine that with your class 0 PR score, your F1-score will be TERRIBLE, which is the correct conclusion for your scenario.\n\nTL,DR: Look at PR scores for ALL your classes, and combine them with a metric like F1-score to have a realistic conclusion about your model performance. The F1-score for your scenario will be TERRIBLE, which is the correct conclusion for your scenario.","metadata":{}},{"cell_type":"markdown","source":"# **Metrics for Multi-Class-Label Classification**\n\nMost Metrics discussed in Binary Classification can be extended to Multi-Class Classification.","metadata":{}},{"cell_type":"markdown","source":"## Multi-Class ROC","metadata":{}},{"cell_type":"markdown","source":"### Intuition\n\nROC is originally used for Binary Classification only, a natural extension to Multi-Class model is the \nIn multi-class model, we can plot N number of AUC ROC Curves for N number classes using One vs ALL methodology. So for Example, If you have three classes named X, Y and Z, you will have one ROC for X classified against Y and Z, another ROC for Y classified against X and Z, and a third one of Z classified against Y and X.\n\n\nFirstly, you need to make use of the below code in [source](https://github.com/scikit-learn/scikit-learn/blob/0fb307bf3/sklearn/metrics/_ranking.py#L690) where we are using the concept of One-Vs-All (ovr) and first thing first, for all `y_true` labels, we need to `label_binarize` them. As we can see, we must pass in the `y_true` and `classes` in which if our classes are `[0,1,2,3,4,5]` then we need to specify in the `labels` argument of `roc_auc_curve`. If we do not specify, the `_encode` will help us as well, so it is up to one's preference if your labels order matter.\n\n```python\nelse:\n    # ovr is same as multi-label\n    y_true_multilabel = label_binarize(y_true, classes=classes)\n    return _average_binary_score(_binary_roc_auc_score, y_true_multilabel,\n                                 y_score, average,\n                                 sample_weight=sample_weight)\n```","metadata":{}},{"cell_type":"markdown","source":"### Implementation of OVR Multi-Class ROC","metadata":{}},{"cell_type":"markdown","source":"#### Step 1: Problem Setup\n\n- Multi-Class: 3 classes of 0, 1 and 2. \n- Number of samples: 4 \n- Predictions: output using Softmax","metadata":{}},{"cell_type":"code","source":"import scipy\n\ny_true_multiclass = np.asarray([0, 1, 2, 1])\n\ny_logit_multiclass = np.asarray(\n    [\n        [0.0802, 0.0347, 0.05],\n        [0.2640, -0.0701, 0.03],\n        [-0.0087, 0.0502, 0.0039],\n        [0.0496, 0.0059, 0.0123],\n    ]\n)\n\ny_pred_multiclass = scipy.special.softmax(y_logit_multiclass, axis=1)\n\ny_pred_multiclass = np.asarray(\n    [\n        [0.34, 0.33, 0.33],\n        [0.4, 0.29, 0.32], \n        [0.33, 0.35, 0.33], \n        [0.34, 0.33, 0.33]]\n)\n","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.308358Z","iopub.execute_input":"2021-09-20T10:37:08.308678Z","iopub.status.idle":"2021-09-20T10:37:08.325055Z","shell.execute_reply.started":"2021-09-20T10:37:08.308629Z","shell.execute_reply":"2021-09-20T10:37:08.323862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 2: Binarize \n\nWe need to binarize the `y_true_multiclass`.\n\n```python\narray([[1, 0, 0],\n       [0, 1, 0],\n       [0, 0, 1],\n       [0, 1, 0]])\n```\n\nwhere we need to interpret as follows:\n\n| Class 1 | Class 2 | Class 3 |\n| --      | --      | --      |\n| 1       | 0       | 0       |\n| 0       | 1       | 0       |\n| 0       | 0       | 1       |\n| 0       | 1       | 0       |\n\nwhere [1,0,0,0] (first column) represents the case where class 1 is the **positive class** and class 2 and 3 are considered the negative class (both are class 0). \n\n| Class 1 Preds | Class 2 Preds| Class 3 Preds |\n| --      | --      | --      |\n| 0.34       | 0.33       | 0.33       |\n| 0.4       | 0.29       | 0.32       |\n| 0.33       | 0.35       | 0.33       |\n| 0.34       | 0.33       | 0.33       |\n\nwhere [0.34, 0.4, 0.33, 0.34] (first column) represents the probability of class 1 being the positive class. And to avoid confusion, the second column [0.33, 0.29, 0.35, 0.33] represents the probability of class 2 being the positive class (class 1).\n\n---\n\nNote that I labelled the above as class 1, 2 and 3 but in our code it is class 0, 1 and 2. This does not affect the ultimate score but just note in case of confusion.","metadata":{}},{"cell_type":"code","source":"y_binarize=sklearn.preprocessing.label_binarize(y_true_multiclass, classes=[0,1,2])\nprint(y_binarize)","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.326157Z","iopub.status.idle":"2021-09-20T10:37:08.326631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 3: ROC score for each class\n\nAt this step, we calculate the ROC score for each class. We have have 3 different scores, one for each class. That is to say, the first score is class 0 vs the rest, where we treated class 0 as the positive class.","metadata":{}},{"cell_type":"code","source":"num_classes = 3\nfpr_dict = dict()\ntpr_dict = dict()\nthreshold_dict = dict()\nroc_auc_dict = dict()\n\nfor label_num in range(num_classes):\n    y_true_for_curr_class = y_binarize[:, label_num]\n    y_pred_for_curr_class = y_pred_multiclass[:, label_num]\n    \n    # calculate fpr,tpr and thresholds across various decision thresholds; pos_label = 1 because one hot encode guarantees it\n    \n    fpr_dict[label_num], tpr_dict[label_num], threshold_dict[label_num] = sklearn.metrics.roc_curve(\n        y_true=y_true_for_curr_class, y_score=y_pred_for_curr_class, pos_label=1\n    )\n    roc_auc_dict[label_num] = sklearn.metrics.auc(fpr_dict[label_num], tpr_dict[label_num])\n    \n    print(f\"ROC score for class {label_num} is {roc_auc_dict[label_num]}. \\nNote we are considering class {label_num} as the positive class and treating other classes as negative class.\\n\")\n\nprint(fpr_dict)\nprint(tpr_dict)\nprint(threshold_dict)","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.32771Z","iopub.status.idle":"2021-09-20T10:37:08.328227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plotting    \nplt.plot(fpr_dict[0], tpr_dict[0], linestyle='--',color='orange', label='Class 0 vs Rest')\nplt.plot(fpr_dict[1], tpr_dict[1], linestyle='--',color='green', label='Class 1 vs Rest')\nplt.plot(fpr_dict[2], tpr_dict[2], linestyle='--',color='blue', label='Class 2 vs Rest')\n\nplt.title('Multiclass ROC curve')\nplt.xlabel('False Positive Rate')\nplt.ylabel('True Positive rate')\nplt.legend(loc='best')\nplt.savefig('Multiclass ROC',dpi=300);    ","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.329076Z","iopub.status.idle":"2021-09-20T10:37:08.329549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 4: One Vs Rest\n\nWe will do a arithmetic mean over the scores we get. This concludes the OvR algorithm.","metadata":{}},{"cell_type":"code","source":"macro_average_roc_score = np.mean(list(roc_auc_dict.values()))\nprint(macro_average_roc_score)","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.330449Z","iopub.status.idle":"2021-09-20T10:37:08.330923Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Step 5: Modularize","metadata":{}},{"cell_type":"code","source":"# This code belows also WORKS for Binary class if you are using Softmax Predictions!\n# replicating from https://scikit-learn.org/stable/auto_examples/model_selection/plot_roc.html#multiclass-settings\ndef multiclass_roc(y_true, y_logit, config):\n    label_dict = dict()\n    fpr = dict()\n    tpr = dict()\n    roc_auc = dict()\n    roc_scores = []\n    for label_num in range(len(config.class_list)):\n\n        # get y_true_multilabel binarized version for each loop (end of each epoch)\n        y_true_multiclass_array = sklearn.preprocessing.label_binarize(\n            y_true, classes=config.class_list\n        )\n\n        y_true_for_curr_class = y_true_multiclass_array[:, label_num]\n        y_pred_for_curr_class = y_logit[:, label_num]\n        # calculate fpr,tpr and thresholds across various decision thresholds\n        # pos_label = 1 because one hot encode guarantees it\n        fpr[label_num], tpr[label_num], _ = sklearn.metrics.roc_curve(\n            y_true=y_true_for_curr_class, y_score=y_pred_for_curr_class, pos_label=1\n        )\n        roc_auc[label_num] = sklearn.metrics.auc(fpr[label_num], tpr[label_num])\n        roc_scores.append(roc_auc[label_num])\n        # if binary class, the one hot encode will (n_samples,1) and therefore will only need to slice [:,0] ONLY.\n        # that is why usually for binary class, we do not need to use this piece of code, just for testing purposes.\n        # However, it will now treat our 0 (negative class) as positive, hence returning the roc for 0, in which case\n        # to get both 0 and 1, you just need to use 1-roc(0)value\n        if config.num_classes == 2:\n            roc_auc[config.class_list[1]] = 1 - roc_auc[label_num]\n            break\n    avg_roc_score = np.mean(roc_scores)\n    return roc_auc, avg_roc_score","metadata":{"execution":{"iopub.status.busy":"2021-09-20T10:37:08.331832Z","iopub.status.idle":"2021-09-20T10:37:08.332314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Multi-Label ROC\n\nIncidentally, the way we compute OVR Multi-Class ROC can be used in Multi-Label ROC as well.","metadata":{}},{"cell_type":"markdown","source":"# Quadratic Weighted Kappa\n\nThe below explanation will correspond to my [notebook during the PANDAS competition](https://www.kaggle.com/reighns/understanding-the-quadratic-weighted-kappa/).","metadata":{}},{"cell_type":"markdown","source":" ## Table of Contents\n1. [Intuition of QWK](#intuition)\n2. [Step 1: Create the NxN histogram matrix O](#confusion) <br>\n3. [Step 2: Create the Weighted Matrix w](#weighted)<br>\n4. [Step 3: Create the Expected Matrix](#expected) <br>\n5. [Step 4: Final Step: Weighted Kappa formula and Its python codes](#qwk) \n\n    ","metadata":{}},{"cell_type":"markdown","source":"## **Intuition of QWK** <a id=\"intuition\"></a> <a id=\"intuition\"></a>\n\nTLDR: one can skip to the last section on the python code implementation of QWK and also take reference to [CPMP's Fast QWK Computation](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145105) as well.\n\nKappa or Cohen’s Kappa is like classification accuracy, except that it is normalized at the baseline of random chance on your dataset: It basically tells you how much better your classifier is performing over the performance of a classifier that simply guesses at random according to the frequency of each class.","metadata":{}},{"cell_type":"markdown","source":"First off, we define the formula exactly as mentioned. Quoting from the evaluation page: **Submissions are scored based on the quadratic weighted kappa, which measures the agreement between two outcomes. This metric typically varies from 0 (random agreement) to 1 (complete agreement). In the event that there is less agreement than expected by chance, the metric may go below 0.** \n\n**The quadratic weighted kappa is calculated as follows. First, an N x N histogram matrix O is constructed, such that $O_{i,j}$ corresponds to the number of `isup_grade`'s i (actual) that received a predicted value j. An N-by-N matrix of weights, w, is calculated based on the difference between actual and predicted values:**\n\n$$w_{i,j} = \\dfrac{(i-j)^2}{(N-1)^2}$$\n\n<br>\n\n**An N-by-N histogram matrix of expected outcomes, E, is calculated assuming that there is** ***no correlation*** **between values.This is calculated as the outer product between the actual histogram vector of outcomes and the predicted histogram vector, normalized such that E and O have the same sum. **\n\n<br>\n\nFinally, from these three matrices, the quadratic weighted kappa is calculated as:\n\n\n$$\\kappa = 1 - \\dfrac{\\sum_{i,j}\\text{w}_{i,j}O_{i,j}}{\\sum_{i,j}\\text{w}_{i,j}E_{i,j}}$$\n\nwhere $w$ is the weighted matrix, $O$ is the histogram matrix and $E$ being the expected matrix.","metadata":{}},{"cell_type":"markdown","source":"## **Step 1: Create the NxN histogram matrix O** <a id=\"confusion\"></a>","metadata":{}},{"cell_type":"markdown","source":"***Warning: This is a counter example to the wrong usage of Quadratic Weighted Kappa. We shall see why soon.***\n\n***Warning: This is a counter example to the wrong usage of Quadratic Weighted Kappa. We shall see why soon.***\n\n***Warning: This is a counter example to the wrong usage of Quadratic Weighted Kappa. We shall see why soon.***\n\n<br>\n\n***Reminder: Although it is a counter example, it still illustrates what a NxN histogram matrix is! We will now call our histogram matrix C instead because in actual fact, the histogram matrix is merely a multi class confusion matrix between actual and predicted values***","metadata":{}},{"cell_type":"markdown","source":"We use a naive example where there are 5 classes (note our competition is_up grade has 6 classes; but this is just an example. Our `y_true` is the ground truth labels and correspondingly, our `y_pred` is the predicted values.","metadata":{}},{"cell_type":"code","source":"y_true = pd.Series(['cat',  'cat', 'dog', 'cat',   'cat',  'cat', 'pig',  'pig', 'hen', 'pig'], name = 'Actual')\ny_pred   = pd.Series(['bird', 'hen', 'pig','bird',  'bird', 'bird', 'pig', 'pig', 'hen', 'pig'], name = 'Predicted')\n\nprint(\"Ground truth:\\n{}\".format(y_true))\nprint(\"-\"*40)\nprint(\"Predicted Values:\\n{}\".format(y_pred))","metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:15:51.378798Z","start_time":"2020-04-22T04:15:51.372814Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.333156Z","iopub.status.idle":"2021-09-20T10:37:08.333823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First, an N x N confusion matrix **C** is constructed, such that $\\text{C}_{i,j}$ is the entry that corresponds to the **number of animal i (actual) that received a predicted value j**. ","metadata":{}},{"cell_type":"code","source":"classes= ['bird','cat','dog','hen', 'pig']\n\n# thank you https://datascience.stackexchange.com/questions/40067/confusion-matrix-three-classes-python \ndef plot_confusion_matrix(cm, classes,\n                          normalize=False,\n                          title='Confusion matrix',\n                          cmap=plt.cm.Blues):\n    \"\"\"\n    This function prints and plots the confusion matrix.\n    Normalization can be applied by setting `normalize=True`.\n    \"\"\"\n    \n    if normalize:\n        cm = cm.astype('float') / cm.sum(axis=1)[:, np.newaxis]\n        print(\"Normalized confusion matrix\")\n    else:\n        print('Confusion matrix, without normalization')\n\n    plt.imshow(cm, interpolation='nearest', cmap=cmap)\n    plt.title(title)\n    plt.colorbar()\n    tick_marks = np.arange(len(classes))\n    plt.xticks(tick_marks, classes, rotation=45)\n    plt.yticks(tick_marks, classes)\n\n    fmt = '.2f' if normalize else 'd'\n    thresh = cm.max() / 2.\n    for i, j in itertools.product(range(cm.shape[0]), range(cm.shape[1])):\n        plt.text(j, i, format(cm[i, j], fmt),\n                 horizontalalignment=\"center\",\n                 color=\"white\" if cm[i, j] > thresh else \"black\")\n\n    plt.ylabel('True label')\n    plt.xlabel('Predicted label')\n    plt.tight_layout()\n    \ncnf_matrix = confusion_matrix(y_true, y_pred,labels=['bird','cat','dog','hen', 'pig'],)\nnp.set_printoptions(precision=2)\n\n# Plot non-normalized confusion matrix\nplt.figure()\nplot_confusion_matrix(cnf_matrix, classes=['bird','cat','dog','hen', 'pig'],\n                      title='Confusion matrix C, without normalization')","metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:17:36.46162Z","start_time":"2020-04-22T04:17:36.288531Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.334628Z","iopub.status.idle":"2021-09-20T10:37:08.335075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example using our competition's dataset","metadata":{}},{"cell_type":"markdown","source":"The above matrix is a **multi class confusion matrix**. As an example, for the 2nd row, we predicted 4 cats to be birds, and 1 cat to be predicted as hen. More compactly, it can be represented as the matrix $C_{2,1}$ = 4.","metadata":{}},{"cell_type":"markdown","source":"We can easily reconcile our example above to relate back to our competition: **After the biopsy is assigned a Gleason score, it is converted into an ISUP grade on a 1-5 scale**. For now, I will include the label **0** because it is the background/non-tissue. \n\nWhat I do next is to take the ground truth and call it `y_true` which is a series. I then generate a dummy `y_pred` by using `np.random.choice` and randomly generate numbers from 0 to 5.","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv(\"../input/prostate-cancer-grade-assessment/train.csv\")\ny_true = train.isup_grade\ny_true\n\ny_pred = np.random.choice(6, 10616, replace=True)\ny_pred = pd.Series(y_pred)\ny_pred","metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:43:40.584317Z","start_time":"2020-04-22T04:43:40.566357Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.335786Z","iopub.status.idle":"2021-09-20T10:37:08.336226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The following confusion matrix, is what we mean by the \"N by N\" (6 by 6) **histogram matrix**. ","metadata":{}},{"cell_type":"code","source":"cm = confusion_matrix(y_true, y_pred,labels=[0,1,2,3,4,5],)\nplot_confusion_matrix(cm, classes=[0,1,2,3,4,5],\n                      title='Confusion matrix C, without normalization')","metadata":{"ExecuteTime":{"end_time":"2020-04-22T04:43:43.243947Z","start_time":"2020-04-22T04:43:42.998584Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.336962Z","iopub.status.idle":"2021-09-20T10:37:08.337358Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***So far we have settled the first portion, construction the histogram matrix.***","metadata":{}},{"cell_type":"markdown","source":"## **Step 2: Create the Weighted Matrix w** <a id =\"weighted\"></a>","metadata":{}},{"cell_type":"markdown","source":"An N-by-N matrix of weights, w, is calculated based on the difference between actual and predicted rating scores. The formula is as follows:  $$w_{i,j} = \\dfrac{(i-j)^2}{(N-1)^2}$$\n\n\nUnder Step-2, each element is weighted. Predictions that are further away from actuals are marked harshly than predictions that are closer to actuals. However, I suddenly realized that these animals above have no obvious hierarchy in them. Therefore, it is not correct to say that predicting a cat as a bird is any better/worse off than predicting a cat as hen.\n\n**Consequently, this is a realization. Because now I start to understand why certain competitions use metrics like quadratic weighted kappa!** Consider the same example as animals just now, but instead of animals, we change to **isup_grade**). So there is an inherent order within the **isup_grade** 0 - 5 such that 0 and 1 is closer than 0 and 2, 1 and 2 is closer to 1 and 3 etc. $$0 > 1 > 2 > 3 > 4 > 5$$\n\n<br>\n\nFor example, let's use a simplified example:\n\n    y_true = [2,2,2,1,2,3,4,5,0,1]\n    y_pred = [1,2,4,1,2,3,4,5,0,1]\n    \nAs a result, our purpose of the weighted matrix is to allocate a higher penalty score if our prediction is further away from the actual value. That is, if our **isup_grade** is 2 but we predicted it as 1 (see the example), then based on our formula above, we have i = 2 and j = 1 (entry $C_{2,1}$), the penalty is $$\\dfrac{(2-1)^2}{(5-1)^2}  = 0.0625$$ but if our **isup_grade** is 2 and we predicted it as 4 (see the example), then the penalty involved is higher $$\\dfrac{(2-4)^2}{(5-1)^2} = 0.25$$\n\nIndeed, the weight matrix helped us assign a heavier penalty to predicting 2 as 4 than 2 as 1.\n\n***Lastly, we also observe that the formula will give us 0 for the diagonals of the weighted matrix. This means penalty is 0 whenever we correctly predict something.***","metadata":{}},{"cell_type":"markdown","source":"To calculate weighted matrix in python code, here is the code with reference to [Aman Arora](https://www.kaggle.com/aroraaman/quadratic-kappa-metric-explained-in-5-simple-steps).","metadata":{}},{"cell_type":"code","source":"# We construct the weighted matrix starting from a zero matrix, it is like constructing a \n# list, we usually start from an empty list and add things inside using loops.\n\ndef weighted_matrix(N):\n    weighted = np.zeros((N,N)) \n    for i in range(len(weighted)):\n        for j in range(len(weighted)):\n            weighted[i][j] = float(((i-j)**2)/(N-1)**2) \n    return weighted\n        \nprint(weighted_matrix(5))","metadata":{"ExecuteTime":{"end_time":"2020-04-22T05:08:16.156024Z","start_time":"2020-04-22T05:08:16.152037Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.338095Z","iopub.status.idle":"2021-09-20T10:37:08.338493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Remember, the further away from the diagonal you get, the worse off is your prediction and it will be penalized harder by a bigger weight.***  As you can easily infer from the weighted matrix above, we use the first row as an example in case one did not understand. Basically, the weighted matrix's first row's first element is 0, because it means we predicted correctly and no penalty is meted out; but as we move further to the left, you can see that the punishment gets harsher and harsher: $$0 < 0.06 < 0.25 < 0.56 < 1$$","metadata":{}},{"cell_type":"markdown","source":"## **Step 3: Create the Expected Matrix** <a id =\"expected\"></a>","metadata":{}},{"cell_type":"code","source":"## dummy example \nactual = pd.Series([2,2,2,3,4,5,5,5,5,5]) \npred   = pd.Series([2,2,2,3,2,1,1,1,1,3]) \nC = confusion_matrix(actual, pred)\n\nN=5\nact_hist=np.zeros([N])\nfor item in actual: \n    act_hist[item - 1]+=1\n    \npred_hist=np.zeros([N])\nfor item in pred: \n    pred_hist[item - 1]+=1\n    \n\nprint(f'Actuals value counts:{act_hist}, \\nPrediction value counts:{pred_hist}')","metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.669296Z","start_time":"2020-04-21T17:19:43.664309Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.339267Z","iopub.status.idle":"2021-09-20T10:37:08.339674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This part is the most difficult to understand, especially for someone who has little statistic backgrounds (mind you when I was majoring in applied math back then, I only took **one** statistic module in my whole tenure.) Googling the idea of expected matrix is not apparently clear to me, but luckily someone pointed to me to read expected frequency of Chi-Square test and then and, there I start to slowly understand.\n\nFor the purpose of better understanding, we call our actual values to be rater A and our prediction model to be rater B. We want to quantify the agreement between rater A and rater B. Here are some terminologies to get hold of first.\n\n- There are a total number of $k = 5$ classes in this example;\n- There are a total number of $n = 10$ observations in this example;\n- Define Y to be the random variable that rater A has chosen (aka our actual classes 1,2,3,4,5);\n- and $\\widehat{Y}$ be the random variable that rater $B$ has chosen (aka our predicted classes 1,2,3,4,5 by rater B)\n- $r_i$ be the i-th entry of the column vector for actual value counts shown above, $c_i$ be the i-th entry of the column vector for prediction value counts shown above.\n<br>\n\nThen the probability of rater A choosing class 2 and rater B choosing class 2 for example, is given by $$P(Y = 2 \\text{ and } \\widehat{Y} = 2)  = P(Y = 2) \\cdot P(\\widehat{Y} = 2) = 30\\% \\times 40\\%  = 12\\%$$\n\nThis is under the assumption that both raters are **independent of each other**. Note that $P(Y = 2) = 30\\%$ because as we see from the actual value counts of rater $A$, there are a total of 3 class 2's and therefore the probability of Y being 2 is just the proportion. Similarly, we calculate $\\widehat{Y}$ the same way.\n\nIn general, the formula of rater A choosing class i and rater B choosing (predicting) class j is given as follows: $$P(Y = i \\text{ and }  \\widehat{Y} = j) =  P(Y = i) \\times P (\\widehat{Y} = j)$$\n\n<br>\n\n\nNow the real question comes: On average, if you have $n$ number of points to predict, how many times (what is the frequency) would you **expect** to see rater A choose class i and rater B choose class j.\n\nTo reiterate, recall that the probability of the actual class being 1 **and (super important word here, it means a joint distribution)** and the predicted class to be 1 as well is $P(Y = 1 \\text{ and }  \\widehat{Y} = 1)$, similarly, the probability of the actual class being 1 and the predicted class to be 2 is $P(Y = 1 \\text{ and }  \\widehat{Y} = 2)$. Generalizing, the probability of the actual class being $i$ and the predicted class being $j$ is the joint probability $$P(Y = i \\text{ and }  \\widehat{Y} = j)$$\n\nSo the intuition lies here: based on the joint probability above, what is the expected number of times (frequency) that $Y = i$ and $\\widehat{Y} = j$ happened (this means Y is i but rater B predict $\\widehat{Y}$ as j) out of 10 times? Easy, just use $n \\times P(Y = i \\text{ and }  \\widehat{Y} = j)$. So for rater B, our prediction model, **by just using theoretical probability**, should have $n \\times P(Y = i \\text{ and }  \\widehat{Y} = j)$ for each $i,j$. But in reality, this may not be the case. Reconcile this idea with the classic coin toss example:\n\n- **Example on coin toss:** Expected frequency is defined as the number of times that we predict an event will occur based on a calculation using theoretical probabilities. You know how a coin has two sides, heads or tails? That means that the probability of the coin landing on any one side is 50% (or 1/2, because it can land on one side out of two possible sides). If you flip a coin 1,000 times, how many times would you expect to get heads? About half the time, or 500 out of the 1,000 flips. To calculate the expected frequency, all we need to do is multiply the total number of tosses (1,000) by the probability of getting a heads (1/2), and we get 1,000 * 1/2 = 500 heads. if event A has prob of p happening, and sample size of n, then the average or expected number of times that A happens is $np$. The expected frequency of heads is 500 out of 1,000 total tosses. The expected frequency is based on our knowledge of probability - we haven't actually done any coin tossing. However, if we had enough time on our hands, we could actually flip a coin 1,000 times to experimentally test our prediction. If we did this, we would be calculating the experimental frequency. The experimental frequency is defined as the number of times that we observe an event to occur when we actually perform an experiment, test, or trial in real life. If you flipped a coin 1,000 times and it landed on heads 479 times, then the experimental frequency of heads is 479.\n\n    But wait - didn't we predict the coin would land on heads 500 times? Why did it actually land on heads only 479 times? Well, the expected frequency was just that - what we expected to happen, based on our knowledge of probability. There are no guarantees when it comes to probability, what we expect to happen might differ from what actually happens. \n\n<br>\n\nSince we know $$E_{2,2} = 10 \\times P(Y = 2 \\text{ and } \\widehat{Y} = 2)  = 10 \\times P(Y = 2) \\cdot P(\\widehat{Y} = 2) = 10 \\times 30\\% \\times 40\\%  = 1.2$$\n\nThis $E_{2,2}$ means that we only expect the rater A to choose 2 and rater B to choose 2 at the same time only 1.2 times out of 10 times. In other words, our predicted model classify 2 as 2 1.2 times. However, from our observations in our matrix C (confusion/histogram matrix), rater A choosing 2 and rater B choosing 2 have a frequency of 3! In other words, our predicted model classified 2 as 2 three times! So we kinda exceeded expectation for this particular configuration.\n\n$$C = \\begin{bmatrix}\n0 & 0 & 0 & 0 & 0\\\\\n0 & 3 & 0 & 0 & 0 \\\\\n0 & 0 & 1 & 0 & 0\\\\\n0 & 1 & 0 & 0 & 0\\\\\n4 & 0 & 1 & 0 & 0\\\\\n\\end{bmatrix}$$\n\nLet me give you one more example, $$E_{5,2} = 10 \\times P(Y = 5 \\text{ and } \\widehat{Y} = 2) = 10 \\times P(Y = 5) \\cdot P(\\widehat{Y} = 2) = 10 \\times 50\\% \\times 40\\%  = 2$$\n\nThis means that we only expect the rater A to choose 5 and rater B to choose 2 at the same time only 2 times out of 10 times! But in reality, our observations say that our rater B classified 5 as 2 zero times! \n\n\nWe calculate $E_{i,j}$ given by the formula: $$E_{i,j} = n \\times P(Y = i \\text{ and }  \\widehat{Y} = j) = n \\times P(Y = i) \\times P (\\widehat{Y} = j) = n \\times \\dfrac{r_i}{n} \\times \\dfrac{c_j}{n}$$\n\n$$E = \\begin{bmatrix}\n0 & 0 & 0 & 0 & 0\\\\\n1.2 & 1.2 & 0.6 & 0 & 0 \\\\\n0.4 & 0.4 & 0.2 & 0 & 0\\\\\n0.4 & 0.4 & 0.2 & 0 & 0\\\\\n2 & 2 & 1 & 0 & 0\\\\\n\\end{bmatrix}$$\n\nNote $r_i \\times c_j$ is the $(i,j)$ entry of the outer product between the actual histogram vector of outcomes (actual value counts) and the predicted histogram vector (prediction value counts).","metadata":{}},{"cell_type":"markdown","source":"What does the random chance mean here? Simple, it just means that the probability for rater A (actual) to be class i AND for rater B (predicted) to be class j is $p\\%$ (say 10%). Therefore, if we have made 100 predictions, random chance aka the theoratical probability tells us you should only have $100 \\times 10\\% = 10$ predictions to be of this configuration (rater A class i AND rater B class j).","metadata":{}},{"cell_type":"markdown","source":"#### Writing out the expected matrix in python\n\nSo to get the expected matrix, E, is calculated assuming that there is no correlation between values.  This is calculated as the outer product between the actual histogram vector of outcomes and the predicted histogram vector, normalized such that E and C have the same sum. Note carefully below that we do not need to normalize both, we just need to normalize E to the same sum as C.","metadata":{}},{"cell_type":"code","source":"E = np.outer(act_hist, pred_hist)/10\n\n\nE\nC","metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.752074Z","start_time":"2020-04-21T17:19:43.74609Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.340434Z","iopub.status.idle":"2021-09-20T10:37:08.340874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Step 4: Final Step: Weighted Kappa formula and Its python codes** <a id=\"qwk\"></a>\n \nFrom these three matrices: E, C and weighted, the quadratic weighted kappa is calculated as: \n\n$$\\kappa = 1 - \\dfrac{\\sum_{i,j}\\text{weighted}_{i,j}C_{i,j}}{\\sum_{i,j}\\text{weighted}_{i,j}E_{i,j}}$$\n\nNote that a higher value generally means your prediction model is way better than a random model, but there is no consensus on which value is really good or bad.","metadata":{}},{"cell_type":"markdown","source":"$$\\text{Weighted} = \\begin{bmatrix}\n0 & 0.0625 & 0.25 & 0.5625 & 1\\\\\n0.0625 & 0 & 0.0625 & 0.25 & 0.5625 \\\\\n0.25 & 0.0625 & 0 & 0.0625 & 0.25\\\\\n0.5625 & 0.25 & 0.0625 & 0 & 0.0625\\\\\n1 & 0.5625 & 0.25 & 0.0625 & 0\\\\\n\\end{bmatrix}$$","metadata":{}},{"cell_type":"markdown","source":"The notation $\\sum_{i,j}\\text{W}_{i,j}C_{i,j}$ is just $$\\sum_{i=1}^{k}\\sum_{j=1}^{k} W_{i,j} C_{i,j} = (W_{1,1}C_{1,1} + W_{1,2}C_{1,2} + ...+ W_{1,k}C_{1,k}) + (W_{2,1}C_{2,1}+...+W_{2,k}C_{2,k}) +...+(W_{k,1}C_{k,1}+...+W_{k,k}C_{k,k})$$\n\nTo put our understanding into perspective, consider just one entry $W_{5,1}C_{5,1} = 1 \\times 4 = 4$. This roughly means that our predicted model classified class 5 as class 1 FOUR times (re: $C_{5,1} = 4$), and since class 5 is so far away from class 1, we need to **punish** this wrong prediction more than the others. And we did see that the corresponding weight $W_{5,1} = 1$ is the highest weight.\n\nConsequently, the numerator being $\\sum_{i,j}\\text{W}_{i,j}C_{i,j}$ calculates the total \"penalty cost\" for the rater A (our predicted model), and similarly, $\\sum_{i,j}\\text{W}_{i,j}E_{i,j}$ calculates the total \"penalty cost\" for the rater B (our \"expected\" model). Therefore, you can think both values (num and den) as a cost function, and the lesser the better. And kappa formula tells us, if our $\\sum_{i,j}\\text{W}_{i,j}C_{i,j}$ is significantly smaller than $\\sum_{i,j}\\text{W}_{i,j}E_{i,j}$, this will yield a very small value of $$\\dfrac{\\sum_{i,j}\\text{weighted}_{i,j}C_{i,j}}{\\sum_{i,j}\\text{weighted}_{i,j}E_{i,j}}$$ which will yield a very high kappa value - signifying a better model.","metadata":{}},{"cell_type":"code","source":"# Method 1\n# apply the weights to the confusion matrix\nweighted = weighted_matrix(5)\nnum = np.sum(np.multiply(weighted, C))\n# apply the weights to the histograms\nden = np.sum(np.multiply(weighted, E))\n\nkappa = 1-np.divide(num,den)\nkappa","metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.798949Z","start_time":"2020-04-21T17:19:43.793962Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.341601Z","iopub.status.idle":"2021-09-20T10:37:08.342034Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Method 2\n\nnum=0\nden=0\nfor i in range(len(weighted)):\n    for j in range(len(weighted)):\n        num+=weighted[i][j]*C[i][j]\n        den+=weighted[i][j]*E[i][j]\n\nweighted_kappa = (1 - (num/den)); weighted_kappa","metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.80593Z","start_time":"2020-04-21T17:19:43.799946Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.342738Z","iopub.status.idle":"2021-09-20T10:37:08.343183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Method 3: Just use sk learn library\n\ncohen_kappa_score(actual, pred, labels=None, weights= 'quadratic', sample_weight=None)","metadata":{"ExecuteTime":{"end_time":"2020-04-21T17:19:43.812911Z","start_time":"2020-04-21T17:19:43.806927Z"},"execution":{"iopub.status.busy":"2021-09-20T10:37:08.343946Z","iopub.status.idle":"2021-09-20T10:37:08.344353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Just a side note that one can also use SK learn's cohen_kappa_score function calculate the quadratic weighted kappa in this competition, with weights set to quadratic. Also please do refer to [CPMP's discussion topic for fast QWK computation](https://www.kaggle.com/c/prostate-cancer-grade-assessment/discussion/145105)","metadata":{}},{"cell_type":"markdown","source":"# References\n\n[Reference I](https://towardsdatascience.com/https-medium-com-abrown004-how-to-ease-the-pain-of-working-with-imbalanced-data-a7f7601f18ba)\n\n[Reference II](https://www.knime.com/blog/correcting-predicted-class-probabilities-in-imbalanced-datasets)\n\n[Reference III](https://towardsdatascience.com/handling-imbalanced-datasets-in-machine-learning-7a0e84220f28)\n\n## **Accuracy**\n- [Indicator Function - Wikipedia](https://en.wikipedia.org/wiki/Indicator_function)\n- [Why is accuracy not the best measure for assessing classification models? - StackExchange](https://stats.stackexchange.com/questions/312780/why-is-accuracy-not-the-best-measure-for-assessing-classification-models)\n- [Example when using accuracy as an outcome measure will lead to a wrong conclusion - StackExchange](https://stats.stackexchange.com/questions/368949/example-when-using-accuracy-as-an-outcome-measure-will-lead-to-a-wrong-conclusio)\n\n\n**Precision-Recall**\n- [Precision-Recall: Estimations of Probabilities - Wikipedia](https://en.wikipedia.org/wiki/Precision_and_recall)\n- [Loss function and Decision Function - StatsExchange](https://stats.stackexchange.com/questions/104988/what-is-the-difference-between-a-loss-function-and-decision-function)\n- [Precision and Recall Tradeoff - Google](https://developers.google.com/machine-learning/crash-course/classification/precision-and-recall)\n\n## **Receiver Operating Characteristic (ROC)**\n\n### Interpretation of ROC\n\n- [Wikipedia has an extensive explanation of the probability behind ROC](https://en.wikipedia.org/wiki/Receiver_operating_characteristic)\n- [Probabilistic Perspective of AUC](https://www.alexejgossmann.com/auc/)\n- [AUC - Insider's Guide to the Theory and Applications](https://sinyi-chou.github.io/classification-auc/)\n- [Safe Handling Instructions for Probabilistic Classification](https://www.youtube.com/watch?v=RXMu96RJj_s).\n- [c-statistics](https://stats.stackexchange.com/questions/193138/roc-curve-drawbacks).\n- [SIIM Melanoma ROC](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173020)\n- [An Introduction to ROC analysis](https://www.sciencedirect.com/science/article/abs/pii/S016786550500303X)\n\n### Pros and Cons of AUROC\n\n- [The Relationship Between Precision-Recall and ROC Curves](chrome-extension://efaidnbmnnnibpcajpcglclefindmkaj/viewer.html?pdfurl=https%3A%2F%2Fwww.biostat.wisc.edu%2F~page%2Frocpr.pdf&clen=137145&chunk=true)\n- [Drawbacks of AUROC](https://stats.stackexchange.com/questions/193138/roc-curve-drawbacks).\n- [ROC vs precision-and-recall curves](https://stats.stackexchange.com/questions/7207/roc-vs-precision-and-recall-curves)\n- [ROC vs Precision-recall curves on imbalanced dataset](https://stats.stackexchange.com/questions/262616/roc-vs-precision-recall-curves-on-imbalanced-dataset)\n- [on why AUC can be misleading](https://stats.stackexchange.com/questions/360017/when-is-an-auc-score-misleadingly-high/360040#360040)\n- [AUC scale and threshold invariant - TDS](https://towardsdatascience.com/an-understandable-guide-to-roc-curves-and-auc-and-why-and-when-to-use-them-92020bc4c5c1)\n\n### Implementation\n\n- [Implementation of ROC - roc-curve-and-auc-from-scratch-in-numpy-visualized - TDS](https://towardsdatascience.com/roc-curve-and-auc-from-scratch-in-numpy-visualized-2612bb9459ab)\n- [Trapezoid Rule](https://ece.uwaterloo.ca/~dwharder/NumericalAnalysis/13Integration/comptrap/complete.html)\n\n\n\n\n\n- [On why thresholds return 2 sometimes](https://stackoverflow.com/questions/52358114/why-is-roc-curve-return-an-additional-value-for-the-thresholds-2-0-for-some-cl)(https://stackoverflow.com/questions/23200518/scikit-learn-roc-curve-why-does-it-return-a-threshold-value-2-some-time)\n- [PR-Curve vs ROC-Curve](https://neptune.ai/blog/f1-score-accuracy-roc-auc-pr-auc#:~:text=ROC%20AUC%20vs%20PR%20AUC&text=What%20is%20different%20however%20is,and%20true%20positive%20rate%20TPR.)\n- https://stackoverflow.com/questions/59666138/sklearn-roc-auc-score-with-multi-class-ovr-should-have-none-average-available\n- https://stackoverflow.com/questions/39685740/calculate-sklearn-roc-auc-score-for-multi-class\n- https://datascience.stackexchange.com/questions/36862/macro-or-micro-average-for-imbalanced-class-problems#:~:text=Micro%2Daverage%20is%20preferable%20if,your%20dataset%20varies%20in%20size.\n- https://www.google.com/search?q=roc_auc_score+multiclass+site:stackoverflow.com&rlz=1C1CHBF_enSG891SG891&sxsrf=ALeKk018tRSfmKgIUw63SPI8dsdkvJgPuw:1608711331403&sa=X&ved=2ahUKEwjg7NDb1OPtAhUXVH0KHVNHCmwQrQIoBHoECAMQBQ&biw=1280&bih=610\n- https://stackoverflow.com/questions/56227246/how-to-calculate-roc-auc-score-having-3-classes\n- https://glassboxmedicine.com/2019/02/23/measuring-performance-auc-auroc/\n- https://stackoverflow.com/questions/56227246/how-to-calculate-roc-auc-score-having-3-classes\n\n---\n\n## **Precision-Recall Curve**\n\n### Interpretation of PR\n\n- [Damage Caused by Classification Accuracy and Other Discontinuous Improper Accuracy Scoring Rules - Frank Harrell](https://www.fharrell.com/post/class-damage/)\n- [Classification vs. Prediction - Frank Harrell](https://www.fharrell.com/post/classification/)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}