|
This article is cited in 4 scientific papers (total in 4 papers)
MODELS OF ECONOMIC AND SOCIAL SYSTEMS
Assessing the validity of clustering of panel data by Monte Carlo methods (using as example the data of the Russian regional economy)
I. L. Kirilyuka, O. V. Sen'kob a Institute of Economics, Russian Academy of Sciences,
32 Nakhimovskii pr., Moscow, 117218, Russia
b Federal Research Center Computer Science and Control, Russian Academy of Sciences,
44/2 Vavilova st., Moscow, 119333, Russia
Abstract:
The paper considers a method for studying panel data based on the use of agglomerative hierarchical clustering — grouping objects based on the similarities and differences in their features into a hierarchy of clusters nested into each other. We used 2 alternative methods for calculating Euclidean distances between objects — the distance between the values averaged over observation interval, and the distance using data for all considered years. Three alternative methods for calculating the distances between clusters were compared. In the first case, the distance between the nearest elements from two clusters is considered to be distance between these clusters, in the second — the average over pairs of elements, in the third — the distance between the most distant elements. The efficiency of using two clustering quality indices, the Dunn and Silhouette index, was studied to select the optimal number of clusters and evaluate the statistical significance of the obtained solutions. The method of assessing statistical reliability of cluster structure consisted in comparing the quality of clustering on a real sample with the quality of clustering on artificially generated samples of panel data with the same number of objects, features and lengths of time series. Generation was made from a fixed probability distribution. At the same time, simulation methods imitating Gaussian white noise and random walk were used. Calculations with the Silhouette index showed that a random walk is characterized not only by spurious regression, but also by “spurious clustering”. Clustering was considered reliable for a given number of selected clusters if the index value on the real sample turned out to be greater than the value of the 95% quantile for artificial data. A set of time series of indicators characterizing production in the regions of the Russian Federation was used as a sample of real data. For these data only Silhouette shows reliable clustering at the level $p<0.05$. Calculations also showed that index values for real data are generally closer to values for random walks than for white noise, but it have significant differences from both. Since three-dimensional feature space is used, the quality of clustering was also evaluated visually. Visually, one can distinguish clusters of points located close to each other, also distinguished as clusters by the applied hierarchical clustering algorithm.
Keywords:
clustering validity, panel data, mesoeconomics, regional economics.
Received: 04.05.2020 Revised: 02.09.2020 Accepted: 18.09.2020
Citation:
I. L. Kirilyuk, O. V. Sen'ko, “Assessing the validity of clustering of panel data by Monte Carlo methods (using as example the data of the Russian regional economy)”, Computer Research and Modeling, 12:6 (2020), 1501–1513; Computer Research and Modeling, 12:6 (2020), e1501–e1513
Linking options:
https://www.mathnet.ru/eng/crm862 https://www.mathnet.ru/eng/crm/v12/i6/p1501
|
Statistics & downloads: |
Abstract page: | 93 | Russian version PDF: | 30 | English version PDF: | 20 | References: | 14 |
|