jsat.clustering.dissimilarity
Class AverageLinkDissimilarity
- java.lang.Object
-
- jsat.clustering.dissimilarity.AbstractClusterDissimilarity
-
- jsat.clustering.dissimilarity.DistanceMetricDissimilarity
-
- jsat.clustering.dissimilarity.LanceWilliamsDissimilarity
-
- jsat.clustering.dissimilarity.AverageLinkDissimilarity
-
- All Implemented Interfaces:
- ClusterDissimilarity, UpdatableClusterDissimilarity
public class AverageLinkDissimilarity extends LanceWilliamsDissimilarity implements UpdatableClusterDissimilarity
Also known as Group-Average Agglomerative Clustering (GAAC) and UPGMA, this measure computer the dissimilarity by summing the distances between all possible data point pairs in the union of the clusters.
-
-
Constructor Summary
Constructors Constructor and Description AverageLinkDissimilarity()Creates a new AverageLinkDissimilarity using theEuclideanDistanceAverageLinkDissimilarity(DistanceMetric dm)Creates a new AverageLinkDissimilarity
-
Method Summary
All Methods Instance Methods Concrete Methods Modifier and Type Method and Description AverageLinkDissimilarityclone()doubledissimilarity(int i, int ni, int j, int nj, double[][] distanceMatrix)Provides the notion of dissimilarity between two sets of points, that may not have the same number of points.doubledissimilarity(int i, int ni, int j, int nj, int k, int nk, double[][] distanceMatrix)Provides the notion of dissimilarity between two sets of points, that may not have the same number of points.doubledissimilarity(java.util.List<DataPoint> a, java.util.List<DataPoint> b)Provides the notion of dissimilarity between two sets of points, that may not have the same number of points.doubledissimilarity(java.util.Set<java.lang.Integer> a, java.util.Set<java.lang.Integer> b, double[][] distanceMatrix)Provides the notion of dissimilarity between two sets of points, that may not have the same number of points.-
Methods inherited from class jsat.clustering.dissimilarity.LanceWilliamsDissimilarity
dissimilarity
-
Methods inherited from class jsat.clustering.dissimilarity.DistanceMetricDissimilarity
distance
-
Methods inherited from class jsat.clustering.dissimilarity.AbstractClusterDissimilarity
createDistanceMatrix, getDistance, setDistance
-
Methods inherited from class java.lang.Object
equals, getClass, hashCode, notify, notifyAll, toString, wait, wait, wait
-
Methods inherited from interface jsat.clustering.dissimilarity.ClusterDissimilarity
distance
-
-
-
-
Constructor Detail
-
AverageLinkDissimilarity
public AverageLinkDissimilarity()
Creates a new AverageLinkDissimilarity using theEuclideanDistance
-
AverageLinkDissimilarity
public AverageLinkDissimilarity(DistanceMetric dm)
Creates a new AverageLinkDissimilarity- Parameters:
dm- the distance measure to use on individual points
-
-
Method Detail
-
clone
public AverageLinkDissimilarity clone()
- Specified by:
clonein interfaceClusterDissimilarity- Specified by:
clonein interfaceUpdatableClusterDissimilarity- Specified by:
clonein classLanceWilliamsDissimilarity
-
dissimilarity
public double dissimilarity(java.util.List<DataPoint> a, java.util.List<DataPoint> b)
Description copied from interface:ClusterDissimilarityProvides the notion of dissimilarity between two sets of points, that may not have the same number of points.- Specified by:
dissimilarityin interfaceClusterDissimilarity- Overrides:
dissimilarityin classLanceWilliamsDissimilarity- Parameters:
a- the first cluster of pointsb- the second cluster of points- Returns:
- a value >= 0 that describes the dissimilarity of the two clusters. The larger the value, the more different the two clusterings are.
-
dissimilarity
public double dissimilarity(java.util.Set<java.lang.Integer> a, java.util.Set<java.lang.Integer> b, double[][] distanceMatrix)Description copied from interface:ClusterDissimilarityProvides the notion of dissimilarity between two sets of points, that may not have the same number of points. This is done using a matrix containing all pairwise distance computations between all points.- Specified by:
dissimilarityin interfaceClusterDissimilarity- Overrides:
dissimilarityin classLanceWilliamsDissimilarity- Parameters:
a- the first set of indices of the original data set that are in a cluster, which map to distanceMatrixb- the second set of indices of the original data set that are in a cluster, which map to distanceMatrixdistanceMatrix- the upper triangual distance matrix as created byAbstractClusterDissimilarity.createDistanceMatrix(jsat.DataSet, jsat.clustering.dissimilarity.ClusterDissimilarity)- Returns:
- a value >= 0 that describes the dissimilarity of the two clusters. The larger the value, the more different the two clusterings are.
-
dissimilarity
public double dissimilarity(int i, int ni, int j, int nj, double[][] distanceMatrix)Description copied from interface:UpdatableClusterDissimilarityProvides the notion of dissimilarity between two sets of points, that may not have the same number of points. This is done using a matrix containing all pairwise distance computations between all points. This distance matrix will then be updated at each iteration and merging, leaving empty space in the matrix. The updates will be done by the clustering algorithm. Implementing this interface indicates that this dissimilarity measure can be accurately computed in an updatable manner that is compatible with a Lance–Williams update.- Specified by:
dissimilarityin interfaceUpdatableClusterDissimilarity- Overrides:
dissimilarityin classLanceWilliamsDissimilarity- Parameters:
i- the index of cluster i's distance in the original data setni- the number of items in the cluster represented by ij- the index of cluster j's distance in the original data setnj- the number of items in the cluster represented by jdistanceMatrix- a distance matrix originally created byAbstractClusterDissimilarity.createDistanceMatrix(jsat.DataSet, jsat.clustering.dissimilarity.ClusterDissimilarity)- Returns:
- a value >= 0 that describes the dissimilarity of the two clusters. The larger the value, the more different the two clusterings are.
-
dissimilarity
public double dissimilarity(int i, int ni, int j, int nj, int k, int nk, double[][] distanceMatrix)Description copied from interface:UpdatableClusterDissimilarityProvides the notion of dissimilarity between two sets of points, that may not have the same number of points. This is done using a matrix containing all pairwise distance computations between all points. This distance matrix will then be updated at each iteration and merging, leaving empty space in the matrix. The updates will be done by the clustering algorithm. Implementing this interface indicates that this dissimilarity measure can be accurately computed in an updatable manner that is compatible with a Lance–Williams update.
This computes the dissimilarity of the union of clusters i and j, (Ci ∪ Cj), with the cluster k. This method is used by other algorithms to perform an update of the distance matrix in an efficient manner.- Specified by:
dissimilarityin interfaceUpdatableClusterDissimilarity- Overrides:
dissimilarityin classLanceWilliamsDissimilarity- Parameters:
i- the index of cluster i's distance in the original data setni- the number of items in the cluster represented by ij- the index of cluster j's distance in the original data setnj- the number of items in the cluster represented by jk- the index of cluster k's distance in the original data setnk- the number of items in the cluster represented by k a distance matrix originally created byAbstractClusterDissimilarity.createDistanceMatrix(jsat.DataSet, jsat.clustering.dissimilarity.ClusterDissimilarity)- Returns:
- a value >= 0 that describes the dissimilarity of the union of two clusters with a third cluster. The larger the value, the more different the resulting clusterings are.
-
-
DataMelt 3.0 © DataMelt by jWork.ORG