Catalano.MachineLearning.Regression.RegressionTrees
Class RegressionTree
- java.lang.Object
-
- Catalano.MachineLearning.Regression.RegressionTrees.RegressionTree
-
- All Implemented Interfaces:
- IRegression, java.io.Serializable, java.lang.Cloneable
public class RegressionTree extends java.lang.Object implements IRegression, java.io.Serializable
Decision tree for regression. A decision tree can be learned by splitting the training set into subsets based on an attribute value test. This process is repeated on each derived subset in a recursive manner called recursive partitioning.Classification and Regression Tree techniques have a number of advantages over many of those alternative techniques.
- Simple to understand and interpret.
- In most cases, the interpretation of results summarized in a tree is very simple. This simplicity is useful not only for purposes of rapid classification of new observations, but can also often yield a much simpler "model" for explaining why observations are classified or predicted in a particular manner.
- Able to handle both numerical and categorical data.
- Other techniques are usually specialized in analyzing datasets that have only one type of variable.
- Tree methods are nonparametric and nonlinear.
- The final results of using tree methods for classification or regression can be summarized in a series of (usually few) logical if-then conditions (tree nodes). Therefore, there is no implicit assumption that the underlying relationships between the predictor variables and the dependent variable are linear, follow some specific non-linear link function, or that they are even monotonic in nature. Thus, tree methods are particularly well suited for data mining tasks, where there is often little a priori knowledge nor any coherent set of theories or predictions regarding which variables are related and how. In those types of data analytics, tree methods can often reveal simple relationships between just a few variables that could have easily gone unnoticed using other analytic techniques.
Some techniques such as bagging, boosting, and random forest use more than one decision tree for their analysis.
- See Also:
GradientTreeBoost,RandomForest, Serialized Form
-
-
Nested Class Summary
Nested Classes Modifier and Type Class and Description static interfaceRegressionTree.NodeOutputAn interface to calculate node output.
-
Constructor Summary
Constructors Constructor and Description RegressionTree()Initialize a new instance of the RegressionTree class.RegressionTree(DecisionVariable[] attributes)Initialize a new instance of the RegressionTree class.RegressionTree(DecisionVariable[] attributes, double[][] x, double[] y, int M, int S, int[][] order, int[] samples)Constructor.RegressionTree(DecisionVariable[] attributes, int J)Initialize a new instance of the RegressionTree class.RegressionTree(DecisionVariable[] attributes, int J, int[][] order, int[] samples, RegressionTree.NodeOutput output)Constructor.RegressionTree(int J)Initialize a new instance of the RegressionTree class.RegressionTree(int numFeatures, int[][] x, double[] y, int J)Constructor.RegressionTree(int numFeatures, int[][] x, double[] y, int J, int[] samples, RegressionTree.NodeOutput output)Constructor.
-
Method Summary
All Methods Instance Methods Concrete Methods Modifier and Type Method and Description IRegressionclone()Clone of the object.double[]getImportance()Returns the variable importance.intgetNumberOfLeafs()Get number of maximum leafs.voidLearn(DatasetRegression dataset)Learn.voidLearn(double[][] input, double[] output)Learn.doublePredict(double[] feature)Predict.doublePredict(int[] feature)voidsetNumberOfLeafs(int J)Set number of maximum leafs.
-
-
-
Constructor Detail
-
RegressionTree
public RegressionTree()
Initialize a new instance of the RegressionTree class.
-
RegressionTree
public RegressionTree(DecisionVariable[] attributes)
Initialize a new instance of the RegressionTree class.- Parameters:
attributes- Attributes.
-
RegressionTree
public RegressionTree(int J)
Initialize a new instance of the RegressionTree class.- Parameters:
J- the maximum number of leaf nodes in the tree.
-
RegressionTree
public RegressionTree(DecisionVariable[] attributes, int J)
Initialize a new instance of the RegressionTree class.- Parameters:
attributes- the attribute properties.J- the maximum number of leaf nodes in the tree.
-
RegressionTree
public RegressionTree(DecisionVariable[] attributes, int J, int[][] order, int[] samples, RegressionTree.NodeOutput output)
Constructor. Learns a regression tree for gradient tree boosting.- Parameters:
attributes- the attribute properties.x- the training instances.y- the response variable.order- the index of training values in ascending order. Note that only numeric attributes need be sorted.J- the maximum number of leaf nodes in the tree.samples- the sample set of instances for stochastic learning. samples[i] should be 0 or 1 to indicate if the instance is used for training.
-
RegressionTree
public RegressionTree(DecisionVariable[] attributes, double[][] x, double[] y, int M, int S, int[][] order, int[] samples)
Constructor. Learns a regression tree for random forest.- Parameters:
attributes- the attribute properties.x- the training instances.y- the response variable.order- the index of training values in ascending order. Note that only numeric attributes need be sorted.M- the number of input variables to pick to split on at each node. It seems that dim/3 give generally good performance, where dim is the number of variables.S- number of instances in a node below which the tree will not split, setting S = 5 generally gives good results.samples- the sample set of instances for stochastic learning. samples[i] is the number of sampling for instance i.
-
RegressionTree
public RegressionTree(int numFeatures, int[][] x, double[] y, int J)Constructor. Learns a regression tree on sparse binary samples.- Parameters:
numFeatures- the number of sparse binary features.x- the training instances of sparse binary features.y- the response variable.J- the maximum number of leaf nodes in the tree.
-
RegressionTree
public RegressionTree(int numFeatures, int[][] x, double[] y, int J, int[] samples, RegressionTree.NodeOutput output)Constructor. Learns a regression tree on sparse binary samples.- Parameters:
numFeatures- the number of sparse binary features.x- the training instances.y- the response variable.J- the maximum number of leaf nodes in the tree.samples- the sample set of instances for stochastic learning. samples[i] should be 0 or 1 to indicate if the instance is used for training.
-
-
Method Detail
-
getNumberOfLeafs
public int getNumberOfLeafs()
Get number of maximum leafs.- Returns:
- Number of maximum leafs.
-
setNumberOfLeafs
public void setNumberOfLeafs(int J)
Set number of maximum leafs.- Parameters:
J- Number of maximum leafs.
-
getImportance
public double[] getImportance()
Returns the variable importance. Every time a split of a node is made on variable the impurity criterion for the two descendent nodes is less than the parent node. Adding up the decreases for each individual variable over the tree gives a simple measure of variable importance.- Returns:
- the variable importance
-
Learn
public void Learn(DatasetRegression dataset)
Description copied from interface:IRegressionLearn.- Specified by:
Learnin interfaceIRegression- Parameters:
dataset- Dataset regression.
-
Learn
public void Learn(double[][] input, double[] output)Description copied from interface:IRegressionLearn.- Specified by:
Learnin interfaceIRegression- Parameters:
input- Input.output- Output.
-
Predict
public double Predict(double[] feature)
Description copied from interface:IRegressionPredict.- Specified by:
Predictin interfaceIRegression- Parameters:
feature- Feature.- Returns:
- Value.
-
Predict
public double Predict(int[] feature)
-
clone
public IRegression clone()
Description copied from interface:IRegressionClone of the object.- Specified by:
clonein interfaceIRegression- Overrides:
clonein classjava.lang.Object- Returns:
- A new copy of the object.
-
-
DataMelt 3.0 © DataMelt by jWork.ORG