Outlier Detection with Principal Components (PC) and Angle-Based Outlier Detection (ABOD): a Comparison

PCA and Angle-Based Outlier Factor (ABOF) are effective techniques for detecting outliers, but they operate on different principles and have distinct strengths.
Author

AE Rodriguez

Published

September 9, 2026

While specific fraud and scam patterns differ across payments channels, a clear theme emerged: financial institutions are seeing more fraud attempts and losses, with criminals increasingly using techniques such as impersonation, social engineering and credential compromises.

Federal Reserve Financial Services Risk Office Survey, April 2026

Outlier detection via unsupervised machine learning is an essential tool for fraud detection. It automatically flags unusual data points that break normal behavioral patterns without requiring pre-labeled examples of fraudulent activity. An outlier can be the earliest sign of fraud.

In this tutorial we compare how well Principal component analysis (PCA) and Angle-Based Outlier Detection (ABOD) perform when identifying outliers.

PCA is an unsupervised algorithm typically used for dimensinality reduction. The various dimensions or factors within the data are reduced to an equal number of Principle Components. The First Principal Component (PC1) represents the maximum of the total variance in the observed variables. Before applying PCA, data needs to be scaled or standardized.

Building upon the distance-based outlier detection algorithms ABOD detects outliers by examining the variance of angles between data points.

We start by generating a synthetic data set with arbitrary inserted outliers; we then examine a conventional data set: the Tax Foundation’s tax competitiveness across the states - to identify the “outlier” states within this tax context.

In the appraisal of algorithm performance we turn to Precision and Recall instead of Accuracy. Accuracy can be misleading in the presence of a small proportion of outliers: an algorithm might achieve high Accuracy by simply predicting the majority class (normal instances) while failing to detect outliers.

Precision and Recall are preferred when treating imbalanced data sets.

Precision measures the proportion of true positive predictions among all positive predictions made by the algorithm. This indicates how many of the detected outliers are actually true outliers.

Recall (or sensitivity) measures the proportion of true positives identified out of all actual positives. This reflects the model’s ability to identify all relevant outliers in the dataset.

Preliminaries

options(digits = 3, scipen = 9999)
remove(list = ls())
graphics.off()

suppressWarnings({
  suppressPackageStartupMessages({
    pacman::p_load(factoextra,
                   FactoMineR,
                   tidyverse,
                   psych,
                   naivebayes,
                   MASS,
                   DescTools,
                   GGally,
                   ggthemes,
                   forecast,
                   mvtnorm,
                   caret,
                   ggrepel,
                   cowplot,
                   abodOutlier,
                   randomForest,
                   janitor
    )
  })
})

Generate Toy Data with Outliers

Normal cluster data

n_normal <- 100
mu <- c(5, 5)
Sigma <- matrix(c(3, 2.5, 2.5, 3), nrow = 2) 

data_matrix <- rmvnorm(n = n_normal, mean = mu, sigma = Sigma)

Generate explicit outliers

outliers <- matrix(c(2, 8, 12, 2), nrow = 2, byrow = TRUE)

df <- rbind.data.frame(data_matrix, outliers) # combine

  colnames(df) <- c("Variable1", "Variable2")

labels <- c(rep("Normal", n_normal), rep("Outlier", 2))

mydf = cbind.data.frame(df, labels)

#cor(df$Variable1, df$Variable2)
plot(mydf$Variable1, df$myVariable2, 
     main = "Synthetic Data",
     xlab = "Variable 1", 
     ylab = "Variable 2",
     col = as.factor(mydf$labels), pch = 19, cex = 1.5)

First Principal Component (PCA)

Principal Component Analysis (PCA) maps a data point onto a lower-dimensional line (the first principal component) that represents the main trend of our data. The first principal components only tracks the direction of maximum variance. Anomalies are to be found in the directions perpendicular to the first principal

We need to calculate the reconstruction error: the distance (squared Euclidean distance) between the original data and its reconstructed version. Reconstructing the point means projecting that lower-dimensional value back into the original multi-dimensional feature space. The reconstruction error is the distance between the original data point and its reconstructed version.

The key here is to know that normal points (inliers) sit close to the principal axis and rebuild with low error. Outliers break the expected correlation structure between features, causing a large gap between the original point and its projection.

mypc = prcomp(mydf[,1:2], 
              #scale = TRUE, 
              center = TRUE)
mydf$pc1 = mypc$x[,1]
pc1_loading <- mypc$rotation[, 1]

# Project the scaled data onto the first principal component
#scaled_data <- scale(df[,1:2], 
 #                    center = mypc$center, 
  #                   scale = mypc$scale)

#projected_pc1 <- scaled_data %*% pc1_loading

# Reconstruct the scaled data using *only* the first component
#reconstructed_scaled <- projected_pc1 %*% t(pc1_loading)

reconstructed_scaled <- mydf$pc1 %*% t(pc1_loading)


#' Calculate Reconstruction Error 
#' (Euclidean distance between original and reconstruction)

scaled_data = scale(mydf[,1:2])

pca_scores <- rowSums((scaled_data - reconstructed_scaled)^2)

plot(sort(pca_scores))

mydf = cbind(mydf, pca_scores) |> dplyr::select(-pc1)

Angle-Based Outlier Detection

myabod = abod(mydf[,1:2], method = "complete")
Done:  10 / 102 
Done:  20 / 102 
Done:  30 / 102 
Done:  40 / 102 
Done:  50 / 102 
Done:  60 / 102 
Done:  70 / 102 
Done:  80 / 102 
Done:  90 / 102 
Done:  100 / 102 
mydf = cbind.data.frame(mydf, myabod)

plot(sort(mydf$myabod))

fj <- classInt::classIntervals(mydf$myabod, n = 2, style = "fisher")

mydf <- mydf |>
  as.data.frame() |>
  mutate(
    fj_class = cut( myabod, fj$brks, include.lowest = TRUE)
  )


mydf$ABODclass = ifelse(mydf$myabod < fj$brks[2], "Outlier","Normal")

table(mydf$ABODclass)

 Normal Outlier 
     67      35 
cm = caret::confusionMatrix(as.factor(mydf$ABODclass ),
                            as.factor(mydf$labels),
                            mode = "prec_recall")

precision_val <- cm$byClass["Precision"]
recall_val    <- cm$byClass["Recall"]

cat("PC1 Precision", precision_val, "\n")
PC1 Precision 1 
cat("P1! Recall", recall_val, "\n")
P1! Recall 0.67 
fj <- classInt::classIntervals(mydf$pca_scores, n = 2, style = "fisher")

mydf <- mydf |>
  as.data.frame() |>
  mutate(
    fj_class = cut( pca_scores, fj$brks, include.lowest = TRUE)
  )


mydf$PCclass = ifelse(mydf$pca_scores > fj$brks[2], "Outlier","Normal")

table(mydf$PCclass)

 Normal Outlier 
     86      16 
cm = caret::confusionMatrix(as.factor(mydf$PCclass ),
                       as.factor(mydf$labels),
                       mode = "prec_recall")

precision_val <- cm$byClass["Precision"]
recall_val    <- cm$byClass["Recall"]

cat("PC1 Precision", precision_val, "\n")
PC1 Precision 1 
cat("P1! Recall", recall_val, "\n")
P1! Recall 0.86 
head(arrange(mydf, desc(pca_scores)))
  Variable1 Variable2  labels pca_scores  myabod  fj_class ABODclass PCclass
1    12.000      2.00 Outlier      16.96 0.00929 (2.83,17]   Outlier Outlier
2     9.424      8.77  Normal       7.43 0.01287 (2.83,17]   Outlier Outlier
3     8.785      8.84  Normal       6.45 0.05957 (2.83,17]   Outlier Outlier
4     0.267      2.36  Normal       6.35 0.00561 (2.83,17]   Outlier Outlier
5     2.000      8.00 Outlier       5.54 0.06142 (2.83,17]   Outlier Outlier
6     8.144      8.34  Normal       4.69 0.15731 (2.83,17]   Outlier Outlier
plot_data <- data.frame(
  orig_x1 = mydf[, 1],
  orig_x2 = mydf[, 2],
  recon_x1 = reconstructed_scaled[, 1],
  recon_x2 = reconstructed_scaled[, 2]
)

# 4. Create the comparison plot
ggplot(plot_data) +
  # Draw original points (Blue)
  geom_point(aes(x = orig_x1, y = orig_x2), 
             color = "blue", size = 2.5) +
  # Draw reconstructed points (Red)
  geom_point(aes(x = recon_x1, y = recon_x2), 
             color = "red", size = 2.5) +
  # Connect original to reconstructed to show error distance (Gray dashed lines)
  geom_segment(aes(x = orig_x1, y = orig_x2, xend = recon_x1, yend = recon_x2), 
               linetype = "dashed", color = "gray40") +
  labs(
    title = "Original vs Reconstructed Data",
    subtitle = "Blue = Original Data | Red = Reconstructed on PC1 Line | Dashed = Error",
    x = "Scaled Feature 1",
    y = "Scaled Feature 2"
  ) +
  theme_minimal()

How to Read the Plot

  • Blue Dots: Your actual data points mapped across the two original features.

  • Red Dots: The shadow or projection of those points forced to lie strictly on the single straight line of PC1.

  • Dashed Lines: The error vector. Longer dashed lines mean higher reconstruction loss for that specific data point.

Dedicated to my mother, Nora, who would have been 93 today.