Documentation of 'Catalano.MachineLearning.Regression.RegressionTrees.Learning.RandomForest' Java class
RandomForest
Catalano.MachineLearning.Regression.RegressionTrees.Learning

Class RandomForest

  • All Implemented Interfaces:
    IRegression, java.io.Serializable, java.lang.Cloneable


    public class RandomForest
    extends java.lang.Object
    implements IRegression, java.io.Serializable
    Random forest for regression. Random forest is an ensemble method that consists of many regression trees and outputs the average of individual trees. The method combines bagging idea and the random selection of features.

    Each tree is constructed using the following algorithm:

    1. If the number of cases in the training set is N, randomly sample N cases with replacement from the original data. This sample will be the training set for growing the tree.
    2. If there are M input variables, a number m << M is specified such that at each node, m variables are selected at random out of the M and the best split on these m is used to split the node. The value of m is held constant during the forest growing.
    3. Each tree is grown to the largest extent possible. There is no pruning.
    The advantages of random forest are:
    • For many data sets, it produces a highly accurate model.
    • It runs efficiently on large data sets.
    • It can handle thousands of input variables without variable deletion.
    • It gives estimates of what variables are important in the classification.
    • It generates an internal unbiased estimate of the generalization error as the forest building progresses.
    • It has an effective method for estimating missing data and maintains accuracy when a large proportion of the data are missing.
    The disadvantages are
    • Random forests are prone to over-fitting for some datasets. This is even more pronounced in noisy classification/regression tasks.
    • For data including categorical variables with different number of levels, random forests are biased in favor of those attributes with more levels. Therefore, the variable importance scores from random forest are not reliable for this type of data.
    See Also:
    Serialized Form
    • Method Summary

      All Methods Instance Methods Concrete Methods 
      Modifier and Type Method and Description
      IRegression clone()
      Clone of the object.
      double error()
      Returns the out-of-bag estimation of RMSE.
      double[] getImportance()
      Returns the variable importance.
      void Learn(DatasetRegression dataset)
      Learn.
      void Learn(double[][] input, double[] output)
      Learn.
      double Predict(double[] x)
      Predict.
      int size()
      Returns the number of trees in the model.
      void trim(int T)
      Trims the tree model set to a smaller size in case of over-fitting.
      • Methods inherited from class java.lang.Object

        equals, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait
    • Constructor Detail

      • RandomForest

        public RandomForest()
        Initialize a new instance of the RandomForest class.
      • RandomForest

        public RandomForest(int T)
        Initialize a new instance of the RandomForest class.
        Parameters:
        T - Number of the trees.
      • RandomForest

        public RandomForest(int T,
                            int M)
        Initialize a new instance of the RandomForest class.
        Parameters:
        T - the number of trees.
        M - the number of input variables to be used to determine the decision at a node of the tree. dim/3 seems to give generally good performance, where dim is the number of variables.
      • RandomForest

        public RandomForest(int T,
                            int M,
                            int S)
        Initialize a new instance of the RandomForest class.
        Parameters:
        T - the number of trees.
        M - the number of input variables to be used to determine the decision at a node of the tree. dim/3 seems to give generally good performance, where dim is the number of variables.
        S - the number of instances in a node below which the tree will not split, setting S = 5 generally gives good results.
      • RandomForest

        public RandomForest(DecisionVariable[] attributes)
        Initialize a new instance of the RandomForest class.
        Parameters:
        attributes - Attributes.
      • RandomForest

        public RandomForest(DecisionVariable[] attributes,
                            int T)
        Initialize a new instance of the RandomForest class.
        Parameters:
        attributes - the attribute properties.
        T - Number of trees.
      • RandomForest

        public RandomForest(DecisionVariable[] attributes,
                            int T,
                            int M)
        Initialize a new instance of the RandomForest class.
        Parameters:
        attributes - the attribute properties.
        T - Number of trees.
        M - the number of input variables to be used to determine the decision at a node of the tree. dim/3 seems to give generally good performance, where dim is the number of variables.
      • RandomForest

        public RandomForest(DecisionVariable[] attributes,
                            int T,
                            int M,
                            int S)
        Constructor. Learns a random forest for regression.
        Parameters:
        attributes - the attribute properties.
        T - the Number of trees.
        M - the number of input variables to be used to determine the decision at a node of the tree. dim/3 seems to give generally good performance, where dim is the number of variables.
        S - the number of instances in a node below which the tree will not split, setting S = 5 generally gives good results.
    • Method Detail

      • error

        public double error()
        Returns the out-of-bag estimation of RMSE. The OOB estimate is quite accurate given that enough trees have been grown. Otherwise the OOB estimate can bias upward.
        Returns:
        the out-of-bag estimation of RMSE
      • getImportance

        public double[] getImportance()
        Returns the variable importance. Every time a split of a node is made on variable the impurity criterion for the two descendent nodes is less than the parent node. Adding up the decreases for each individual variable over all trees in the forest gives a fast measure of variable importance that is often very consistent with the permutation importance measure.
        Returns:
        the variable importance
      • size

        public int size()
        Returns the number of trees in the model.
        Returns:
        the number of trees in the model
      • trim

        public void trim(int T)
        Trims the tree model set to a smaller size in case of over-fitting. Or if extra decision trees in the model don't improve the performance, we may remove them to reduce the model size and also improve the speed of prediction.
        Parameters:
        T - the new (smaller) size of tree model set.
      • Learn

        public void Learn(double[][] input,
                          double[] output)
        Description copied from interface: IRegression
        Learn.
        Specified by:
        Learn in interface IRegression
        Parameters:
        input - Input.
        output - Output.
      • Predict

        public double Predict(double[] x)
        Description copied from interface: IRegression
        Predict.
        Specified by:
        Predict in interface IRegression
        Parameters:
        x - Feature.
        Returns:
        Value.
      • clone

        public IRegression clone()
        Description copied from interface: IRegression
        Clone of the object.
        Specified by:
        clone in interface IRegression
        Overrides:
        clone in class java.lang.Object
        Returns:
        A new copy of the object.

DataMelt 3.0 © DataMelt by jWork.ORG

Ads help maintain this website.