| Literature DB >> 35945420 |
Bogdan Kirillov1,2, Maxim Panov3.
Abstract
In this paper we explore the use of income inequality metrics such as Gini or Palma coefficients as a tool to identify anomalies via capsule networks. We demonstrate how the interplay between primary and class capsules gives rise to differences in behavior regarding anomalous and normal input which can be exploited to detect anomalies. Our setup for anomaly detection requires supervision in a form of known outliers. We derive several criteria for capsule networks and apply them to a number of Computer Vision benchmark datasets (MNIST, Fashion-MNIST, Kuzushiji-MNIST and CIFAR10), as well as to the dataset of skin lesion images (HAM10000) and the dataset of CRISPR-Cas9 off-target pairs. The proposed methods outperform the competitors in the majority of considered cases.Entities:
Mesh:
Year: 2022 PMID: 35945420 PMCID: PMC9363486 DOI: 10.1038/s41598-022-17734-7
Source DB: PubMed Journal: Sci Rep ISSN: 2045-2322 Impact factor: 4.996
Figure 1Capsule Networks architecture. Primary capsules are formed out of a convolutional layer set. Those layers have the same input and their output is combined for routing. The decoder is not shown. Architecture for MNIST-like experiments is shown here and the changes we made for different cases are described in the respective sections.
Figure 2An example of coupling coefficients distributions for the data from KuzushijiMNIST. The histograms of [Eq. (2)] for normal and anomalous capsules in normal and anomalous cases are shown with specified Gini and Palma coefficients. Note the difference between the distributions for normal (a) and anomalous (b) capsules in the normal case and the anomalous one. Normal capsule only mildly captures the difference between the samples. We use the couplings of anomalous capsule to compute Gini and Palma coefficients to capture this more pronounced difference for anomalous capsule. Note that on the plot, the distribution for the logarithm of the summed couplings are shown, but the values of Gini and Palma coefficients given are computed on the summed couplings without the logarithm.
Detailed information on classes for HAM10000 dataset.
| Skin lesion type | # Images | Is malignant? |
|---|---|---|
| Melanoma | 1113 | Yes |
| Basal cell carcinoma | 514 | Yes |
| Dermatofibroma | 115 | No |
| Melanocytic nevus | 6705 | No |
| Vascular lesion | 142 | No |
| Benign keratosis-like | 1099 | No |
| Intraepithelial carcinoma | 327 | Yes |
AUROCs for diverse outlier setup (fractions 0.1%, 1% and 10%).
| 0.1%, mean ± std | 1%, mean ± std | 10%, mean ± std | Dataset | |
|---|---|---|---|---|
| Palma | 0.7213 ± 0.0644 | CIFAR10 | ||
| Gini | ||||
| Plain | 0.5000 ± 0.0000 | 0.5003 ± 0.0009 | 0.6372 ± 0.0635 | |
| A | 0.5813 ± 0.1159 | 0.5856 ± 0.1207 | 0.6084 ± 0.0985 | |
| 0.6487 ± 0.0755 | 0.7915 ± 0.0581 | |||
| 0.5851 ± 0.1085 | 0.5908 ± 0.1198 | 0.6159 ± 0.0944 | ||
| ABC | 0.6407 ± 0.0714 | 0.6426 ± 0.0772 | 0.6489 ± 0.0707 | |
| NL | 0.5000 ± 0.0811 | 0.5000 ± 0.0811 | 0.5000 ± 0.0812 | |
| Palma | 0.8387 ± 0.0876 | 0.9795 ± 0.0123 | 0.9805 ± 0.0221 | MNIST |
| Gini | 0.8366 ± 0.0883 | |||
| Plain | 0.5767 ± 0.0660 | 0.8633 ± 0.0584 | 0.9821 ± 0.0060 | |
| A | 0.9545 ± 0.0313 | 0.9541 ± 0.0392 | ||
| 0.9144 ± 0.0644 | 0.9528 ± 0.0297 | 0.7479 ± 0.0929 | ||
| ABC | 0.8169 ± 0.1073 | 0.8102 ± 0.1051 | 0.8293 ± 0.1063 | |
| NL | 0.4943 ± 0.1723 | 0.4947 ± 0.1714 | 0.4942 ± 0.1716 | |
| Palma | 0.7458 ± 0.0489 | KMNIST | ||
| Gini | 0.7446 ± 0.0494 | 0.9096 ± 0.0301 | ||
| Plain | 0.5048 ± 0.0098 | 0.6642 ± 0.0692 | 0.9071 ± 0.0362 | |
| A | 0.7585 ± 0.0764 | 0.7881 ± 0.0572 | 0.7870 ± 0.0764 | |
| 0.7931 ± 0.0548 | ||||
| 0.8923 ± 0.0461 | 0.9066 ± 0.0391 | |||
| ABC | 0.6025 ± 0.1122 | 0.6049 ± 0.1127 | 0.6180 ± 0.1215 | |
| NL | 0.5000 ± 0.1078 | 0.5000 ± 0.1072 | 0.5000 ± 0.1067 | |
| Palma | 0.8388 ± 0.1134 | FMNIST | ||
| Gini | 0.8419 ± 0.1119 | |||
| Plain | 0.5759 ± 0.1172 | 0.8027 ± 0.0797 | 0.9210 ± 0.0554 | |
| A | 0.9090 ± 0.0570 | 0.9087 ± 0.0820 | ||
| 0.8792 ± 0.0567 | 0.8834 ± 0.0836 | 0.7453 ± 0.1117 | ||
| 0.9195 ± 0.0715 | 0.9080 ± 0.0875 | |||
| ABC | 0.7921 ± 0.1595 | 0.7921 ± 0.1588 | 0.7933 ± 0.1611 | |
| NL | 0.5000 ± 0.2215 | 0.5000 ± 0.2218 | 0.5000 ± 0.2207 |
The best and the second-best results are in [bold].
AUROCs for diverse inlier setup (fractions 0.1%, 1% and 10%).
| 0.1% mean ± std | 1% mean ± std | 10% mean ± std | Dataset | |
|---|---|---|---|---|
| Palma | 0.7980 ± 0.0624 | CIFAR10 | ||
| Gini | 0.6982 ± 0.0575 | |||
| Plain | 0.5000 ± 0.0000 | 0.5131 ± 0.0157 | 0.7060 ± 0.0768 | |
| A | 0.5111 ± 0.1465 | 0.5217 ± 0.1388 | 0.6329 ± 0.1361 | |
| 0.8341 ± 0.0284 | ||||
| 0.5113 ± 0.1419 | 0.5223 ± 0.1347 | 0.6397 ± 0.1250 | ||
| ABC | 0.5230 ± 0.118 | 0.5335 ± 0.1163 | 0.5195 ± 0.1051 | |
| NL | 0.5000 ± 0.0812 | 0.5000 ± 0.0812 | 0.5000 ± 0.0811 | |
| Palma | 0.9697 ± 0.0035 | MNIST | ||
| Gini | ||||
| Plain | 0.8039 ± 0.0727 | 0.9634 ± 0.0138 | 0.9928 ± 0.0031 | |
| A | 0.9508 ± 0.0342 | 0.9898 ± 0.0080 | 0.9819 ± 0.0442 | |
| 0.9496 ± 0.0568 | 0.6210 ± 0.1288 | 0.3073 ± 0.0997 | ||
| 0.9593 ± 0.0248 | 0.9920 ± 0.0052 | |||
| ABC | 0.5420 ± 0.2011 | 0.5398 ± 0.2025 | 0.5598 ± 0.1960 | |
| NL | 0.5048 ± 0.1731 | 0.5028 ± 0.1722 | 0.5061 ± 0.1723 | |
| Palma | KMNIST | |||
| Gini | 0.9066 ± 0.0285 | |||
| Plain | 0.6029 ± 0.0322 | 0.8303 ± 0.0288 | 0.9558 ± 0.0159 | |
| A | 0.7611 ± 0.0875 | 0.8908 ± 0.0389 | 0.9501 ± 0.0294 | |
| 0.8451 ± 0.0402 | 0.5218 ± 0.0594 | |||
| 0.8147 ± 0.0670 | 0.9286 ± 0.0242 | 0.9784 ± 0.0119 | ||
| ABC | 0.5000 ± 0.1279 | 0.5000 ± 0.1279 | 0.5000 ± 0.1279 | |
| NL | 0.5000 ± 0.1071 | 0.5000 ± 0.1074 | 0.5000 ± 0.1077 | |
| Palma | FMNIST | |||
| Gini | ||||
| Plain | 0.6517 ± 0.1410 | 0.8076 ± 0.1266 | 0.9376 ± 0.0525 | |
| A | 0.7474 ± 0.1705 | 0.8192 ± 0.1257 | 0.8527 ± 0.0923 | |
| 0.8762 ± 0.1262 | 0.7358 ± 0.2231 | 0.5970 ± 0.2339 | ||
| 0.7374 ± 0.2045 | 0.7989 ± 0.1598 | 0.8427 ± 0.1199 | ||
| ABC | 0.5515 ± 0.2192 | 0.5499 ± 0.2133 | 0.5775 ± 0.2011 | |
| NL | 0.5000 ± 0.2215 | 0.5000 ± 0.2208 | 0.5000 ± 0.2214 |
The best and the second-best results are in [bold].
AUROCs for HAM10000 with the following setups: A—diverse outliers, diverse inliers, B—diverse outliers, homogeneous inliers, C—homogeneous outliers, homogeneous inliers, D—homogeneous outliers, diverse inliers.
| A | B | C | D | |
|---|---|---|---|---|
| Palma | ||||
| Gini | ||||
| Plain | 0.5000 ± 0.0000 | 0.5000 ± 0.0000 | 0.4996 ± 0.0005 | 0.5000 ± 0.0000 |
| A | 0.5693 ± 0.0120 | 0.5815 ± 0.0216 | 0.5846 ± 0.0129 | 0.5953 ± 0.0247 |
| 0.6950 ± 0.0055 | 0.7204 ± 0.0120 | 0.7468 ± 0.0099 | 0.7406 ± 0.0152 | |
| 0.5675 ± 0.0122 | 0.5945 ± 0.0200 | 0.5848 ± 0.0132 | 0.6079 ± 0.0234 | |
| ABC | 0.5456 ± 0.0027 | 0.5888 ± 0.0300 | 0.5710 ± 0.0076 | 0.6066 ± 0.0016 |
| NL | 0.5865 ± 0.0030 | 0.5908 ± 0.0046 | 0.6173 ± 0.0037 | 0.6269 ± 0.005 |
The best and the second-best results are in [bold].