jsat
Class DataSet<Type extends DataSet>
- java.lang.Object
-
- jsat.DataSet<Type>
-
- Direct Known Subclasses:
- ClassificationDataSet, RegressionDataSet, SimpleDataSet
public abstract class DataSet<Type extends DataSet> extends java.lang.ObjectThis is the base class for representing a data set. A data set contains multiple samples, each of which should have the same number of attributes. Conceptually, eachDataPointrepresents a row in the data set, and the attributes form the columns.
-
-
Constructor Summary
Constructors Constructor and Description DataSet()
-
Method Summary
All Methods Instance Methods Abstract Methods Concrete Methods Modifier and Type Method and Description voidapplyTransform(DataTransform dt)Applies the given transformation to all points in this data set, replacing each data point with the new value.voidapplyTransform(DataTransform dt, boolean mutate)Applies the given transformation to all points in this data set.voidapplyTransform(DataTransform dt, boolean mutate, java.util.concurrent.ExecutorService ex)Applies the given transformation to all points in this data set in parallel.voidapplyTransform(DataTransform dt, java.util.concurrent.ExecutorService ex)Applies the given transformation to all points in this data set in parallel, replacing each data point with the new value.longcountMissingValues()java.util.List<Type>cvSet(int folds)Creates folds data sets that contain data from this data set.java.util.List<Type>cvSet(int folds, java.util.Random rand)Creates folds data sets that contain data from this data set.CategoricalData[]getCategories()Returns the array containing the categorical data information for this data set.java.lang.StringgetCategoryName(int i)Returns the name used for the i'th categorical attribute.Vec[]getColumnMeanVariance()Computes the weighted mean and variance for each column of feature values.MatrixgetDataMatrix()Creates a matrix from the data set, where each row represent a data point, and each column is one of the numeric example from the data set.MatrixgetDataMatrixView()Creates a matrix backed by the data set, where each row is a data point from the dataset, and each column is one of the numeric examples from the data set.abstract DataPointgetDataPoint(int i)Returns the i'th data point in this set.java.util.Iterator<DataPoint>getDataPointIterator()Returns an iterator that will iterate over all data points in the set.java.util.List<DataPoint>getDataPoints()Creates a list containing the same DataPoints in this set.java.util.List<Vec>getDataVectors()Creates a list of the vectors values for each data point in the correct order.VecgetDataWeights()This method returns the weight of each data point in a single Vector.TypegetMissingDropped()This method returns a dataset that is a subset of this dataset, where only the rows that have no missing values are kept.intgetNumCategoricalVars()Returns the number of categorical variables for each data point in the setVecgetNumericColumn(int i)The data set can be seen as a NxM matrix, were each row is a data point, and each column the values for a particular variable.Vec[]getNumericColumns()Creates an array of column vectors for every numeric variable in this data set.Vec[]getNumericColumns(java.util.Set<java.lang.Integer> skipColumns)Creates an array of column vectors for every numeric variable in this data set.java.lang.StringgetNumericName(int i)Returns the name used for the i'th numeric attribute.intgetNumFeatures()Returns the number of features in this data set, which is the sum ofgetNumCategoricalVars()andgetNumNumericalVars()intgetNumNumericalVars()Returns the number of numerical variables for each data point in the setOnLineStatistics[]getOnlineColumnStats(boolean useWeights)Returns summary statistics computed in an online fashion for each numeric variable.OnLineStatisticsgetOnlineDenseStats()Returns anOnLineStatisticsobject that is built by observing what proportion of each data point contains non zero numerical values.abstract intgetSampleSize()Returns the number of data points in this data setOnLineStatisticsgetSparsityStats()Returns statistics on the sparsity of the vectors in this data set.DataSetgetTwiceShallowClone()Returns a new version of this data set that is of the same type, and contains a different listing pointing to shallow data point copies.java.util.List<Type>randomSplit(double... splits)Splits the dataset randomly into proportionally sized partitions.java.util.List<Type>randomSplit(java.util.Random rand, double... splits)Splits the dataset randomly into proportionally sized partitions.voidreplaceNumericFeatures(java.util.List<Vec> newNumericFeatures)This method will replace every numeric feature in this dataset with a Vec object from the given list.abstract voidsetDataPoint(int i, DataPoint dp)Replaces an already existing data point with the one given.booleansetNumericName(java.lang.String name, int i)Sets the unique name associated with the i'th numeric attribute.abstract DataSet<Type>shallowClone()Returns a new version of this data set that is of the same type, and contains a different list pointing to the same data points.
-
-
-
Method Detail
-
setNumericName
public boolean setNumericName(java.lang.String name, int i)Sets the unique name associated with the i'th numeric attribute. All strings will be converted to lower case first.- Parameters:
name- the name to usei- the ith attribute.- Returns:
- true if the value was set, false if it was not set because an invalid index was given .
-
getNumericName
public java.lang.String getNumericName(int i)
Returns the name used for the i'th numeric attribute.- Parameters:
i- the ith attribute.- Returns:
- the name used for the i'th numeric attribute.
-
getCategoryName
public java.lang.String getCategoryName(int i)
Returns the name used for the i'th categorical attribute.- Parameters:
i- the ith attribute.- Returns:
- the name used for the i'th categorical attribute.
-
applyTransform
public void applyTransform(DataTransform dt)
Applies the given transformation to all points in this data set, replacing each data point with the new value. No mutation of the data points will occur- Parameters:
dt- the transformation to apply
-
applyTransform
public void applyTransform(DataTransform dt, java.util.concurrent.ExecutorService ex)
Applies the given transformation to all points in this data set in parallel, replacing each data point with the new value. No mutation of the data points will occur.- Parameters:
dt- the transformation to applyex- the threadpool to provide threads from. May benullto perform operations in serial
-
applyTransform
public void applyTransform(DataTransform dt, boolean mutate)
Applies the given transformation to all points in this data set. If the transform supports mutating the original data points, this will be applied ifmutableTransformis set totrue- Parameters:
dt- the transformation to applymutate-trueto mutableTransform the original data points,falseto ignore the ability to mutableTransform and replace the original data points.
-
applyTransform
public void applyTransform(DataTransform dt, boolean mutate, java.util.concurrent.ExecutorService ex)
Applies the given transformation to all points in this data set in parallel. If the transform supports mutating the original data points, this will be applied ifmutableTransformis set totrue- Parameters:
dt- the transformation to applymutate-trueto mutableTransform the original data points,falseto ignore the ability to mutableTransform and replace the originalex- the threadpool to provide threads from. May benullto perform operations in serial
-
replaceNumericFeatures
public void replaceNumericFeatures(java.util.List<Vec> newNumericFeatures)
This method will replace every numeric feature in this dataset with a Vec object from the given list. All vecs in the given list must be of the same size.- Parameters:
newNumericFeatures- the list of new numeric features to use
-
getDataPoint
public abstract DataPoint getDataPoint(int i)
Returns the i'th data point in this set. The order will never chance so long as no data points are added or removed from the set.- Parameters:
i- the i'th data point in this set- Returns:
- the i'th data point in this set
-
setDataPoint
public abstract void setDataPoint(int i, DataPoint dp)Replaces an already existing data point with the one given. Any values associated with the data point, but not apart of it, will remain intact.- Parameters:
i- the i'th dataPoint to set.dp- the data point to set at the specified index
-
getOnlineColumnStats
public OnLineStatistics[] getOnlineColumnStats(boolean useWeights)
Returns summary statistics computed in an online fashion for each numeric variable. This returns all summary statistics, but can be less numerically stable and uses more memory.
NaNs / missing values will be ignored in the statistics for each column.- Parameters:
useWeights-trueto return the weighted statistics, unweighted otherwise.- Returns:
- an array of summary statistics
-
getOnlineDenseStats
public OnLineStatistics getOnlineDenseStats()
Returns anOnLineStatisticsobject that is built by observing what proportion of each data point contains non zero numerical values. A mean of 1 indicates all values were fully dense, and a mean of 0 indicates all values were completely sparse (all zeros).- Returns:
- statistics on the percent sparseness of each data point
-
getColumnMeanVariance
public Vec[] getColumnMeanVariance()
Computes the weighted mean and variance for each column of feature values. This has less overhead thangetOnlineColumnStats(boolean)but returns less information.- Returns:
- an array of the vectors containing the mean and variance for each column.
-
getDataPointIterator
public java.util.Iterator<DataPoint> getDataPointIterator()
Returns an iterator that will iterate over all data points in the set. The behavior is not defined if one attempts to modify the data set while being iterated.- Returns:
- an iterator for the data points
-
getSampleSize
public abstract int getSampleSize()
Returns the number of data points in this data set- Returns:
- the number of data points in this data set
-
getNumCategoricalVars
public int getNumCategoricalVars()
Returns the number of categorical variables for each data point in the set- Returns:
- the number of categorical variables for each data point in the set
-
getNumNumericalVars
public int getNumNumericalVars()
Returns the number of numerical variables for each data point in the set- Returns:
- the number of numerical variables for each data point in the set
-
getCategories
public CategoricalData[] getCategories()
Returns the array containing the categorical data information for this data set. Changes to this will be reflected in the data set.- Returns:
- the array of
CategoricalData
-
getMissingDropped
public Type getMissingDropped()
This method returns a dataset that is a subset of this dataset, where only the rows that have no missing values are kept. The new dataset is backed by this dataset.- Returns:
- a subset of this dataset that has all data points with missing features dropped
-
randomSplit
public java.util.List<Type> randomSplit(java.util.Random rand, double... splits)
Splits the dataset randomly into proportionally sized partitions.- Parameters:
rand- the source of randomness for moving data aroundsplits- any array, where the length is the number of datasets to create and the value of in each index is the fraction of samples that should be placed into that dataset. The sum of values must be less than or equal to 1.0- Returns:
- a list of new datasets
-
randomSplit
public java.util.List<Type> randomSplit(double... splits)
Splits the dataset randomly into proportionally sized partitions.- Parameters:
splits- any array, where the length is the number of datasets to create and the value of in each index is the fraction of samples that should be placed into that dataset. The sum of values must be less than or equal to 1.0- Returns:
- a list of new datasets
-
cvSet
public java.util.List<Type> cvSet(int folds, java.util.Random rand)
Creates folds data sets that contain data from this data set. The data points in each set will be random. These are meant for cross validation- Parameters:
folds- the number of cross validation sets to create. Should be greater then 1rand- the source of randomness- Returns:
- the list of data sets.
-
cvSet
public java.util.List<Type> cvSet(int folds)
Creates folds data sets that contain data from this data set. The data points in each set will be random. These are meant for cross validation- Parameters:
folds- the number of cross validation sets to create. Should be greater then 1- Returns:
- the list of data sets.
-
getDataPoints
public java.util.List<DataPoint> getDataPoints()
Creates a list containing the same DataPoints in this set. They are soft copies, in the same order as this data set. However, altering this list will have no effect on DataSet. Altering the DataPoints in the list will effect the DataPoints in this DataSet.- Returns:
- a list of the DataPoints in this DataSet.
-
getDataVectors
public java.util.List<Vec> getDataVectors()
Creates a list of the vectors values for each data point in the correct order.- Returns:
- a list of the vectors for the data points
-
getNumericColumn
public Vec getNumericColumn(int i)
The data set can be seen as a NxM matrix, were each row is a data point, and each column the values for a particular variable. This method grabs all the numerical values for a 'column' and returns it as one vector.
This vector can be altered and will not effect any of the values in the data set- Parameters:
i- the i'th numerical variable to obtain all values of- Returns:
- a Vector of length
getSampleSize()
-
countMissingValues
public long countMissingValues()
- Returns:
- the number of missing values in both numeric and categorical features
-
getNumericColumns
public Vec[] getNumericColumns()
Creates an array of column vectors for every numeric variable in this data set. The index of the array corresponds to the numeric feature index. This method is faster and more efficient than callinggetNumericColumn(int)when multiple columns are needed.
Note, that the columns returned by this method may be cached and re used by the DataSet itself. If you need to alter the columns you should create your own copy of these vectors. If you know that you will be the only person getting a column vector from this data set, then you may safely alter the columns without mutating the data points themselves. However, future callers may or may not receive the same vector objects.- Returns:
- an array of the column vectors
-
getNumericColumns
public Vec[] getNumericColumns(java.util.Set<java.lang.Integer> skipColumns)
Creates an array of column vectors for every numeric variable in this data set. The index of the array corresponds to the numeric feature index. This method is faster and more efficient than callinggetNumericColumn(int)when multiple columns are needed.
A set of columns to skip can be provided in order to save memory if one does not need all the columns.
Note, that the columns returned by this method may be cached and re used by the DataSet itself. If you need to alter the columns you should create your own copy of these vectors. If you know that you will be the only person getting a column vector from this data set, then you may safely alter the columns without mutating the data points themselves. However, future callers may or may not receive the same vector objects.- Parameters:
skipColumns- if a column's index is in this set, anullwill be returned in the array at the column's index instead of a vector- Returns:
- an array of the column vectors
-
getDataMatrix
public Matrix getDataMatrix()
Creates a matrix from the data set, where each row represent a data point, and each column is one of the numeric example from the data set.
This matrix can be altered and will not effect any of the values in the data set.- Returns:
- a matrix of the data points.
-
getDataMatrixView
public Matrix getDataMatrixView()
Creates a matrix backed by the data set, where each row is a data point from the dataset, and each column is one of the numeric examples from the data set.
Any modifications to this matrix will be reflected in the dataset.
This method has the advantage overgetDataMatrix()in that it does not use any additional memory and it maintains any sparsity information.- Returns:
- a matrix representation of the data points
-
getNumFeatures
public int getNumFeatures()
Returns the number of features in this data set, which is the sum ofgetNumCategoricalVars()andgetNumNumericalVars()- Returns:
- the total number of features in this data set
-
shallowClone
public abstract DataSet<Type> shallowClone()
Returns a new version of this data set that is of the same type, and contains a different list pointing to the same data points.- Returns:
- a shallow copy of this data set
-
getTwiceShallowClone
public DataSet getTwiceShallowClone()
Returns a new version of this data set that is of the same type, and contains a different listing pointing to shallow data point copies. Because the data point object contains the weight itself, the weight is not shared - while the vector and array information is. This allows altering the weights of the data points while preserving the original weights.
Altering the list or weights of the returned data set will not be reflected in the original. Altering the feature values will.- Returns:
- a shallow copy of shallow data point copies for this data set.
-
getSparsityStats
public OnLineStatistics getSparsityStats()
Returns statistics on the sparsity of the vectors in this data set. Vectors that are not considered sparse will be treated as completely dense, even if zero values exist in the data.- Returns:
- an object containing the statistics of the vector sparsity
-
getDataWeights
public Vec getDataWeights()
This method returns the weight of each data point in a single Vector. When all data points have the same weight, this will return a vector that uses fixed memory instead of allocating a full double backed array.- Returns:
- a vector that will return the weight for each data point with the same corresponding index.
-
-
DataMelt 3.0 © DataMelt by jWork.ORG