The FOREST Procedure

Isolation Forests

An isolation forest (Liu, Ting, and Zhou (2008)) is a specially constructed forest that is used for anomaly detection instead of target prediction. When the FOREST procedure creates an isolation forest, it outputs anomaly scores in the scored data table that is specified in the OUTPUT statement.

For each split in an isolation forest, one input variable is randomly chosen. If the variable is an interval variable, then it is split at a random value between the maximum and minimum values of the observations in that node. If the variable is a nominal variable, then each level of the variable is assigned to a random branch. By constructing the forest this way, anomalous observations are likely to have a shorter path from the root node to the leaf node than nonanomalous observations have.

The anomaly score, s left-parenthesis x right-parenthesis, of observation x is calculated as:

s left-parenthesis x right-parenthesis equals 2 Superscript minus h left-parenthesis x right-parenthesis

where h left-parenthesis x right-parenthesis is the average, over all trees, of the length of the path from the root node to the leaf node that contains observation x, divided by the average length of all paths across all trees.

The anomaly score is always between 0 and 1, where values closer to 1 indicate a higher chance of the observation being an anomaly.

Last updated: December 09, 2022