SISSA DIGITAL LIBRARYInstitutional Research Information System (Statistiche: prodotti, OA)
Per informazioni contatta sdl@sissa.it

Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a manifold whose Intrinsic Dimension (ID) is much lower than the crude large number of coordinates. Such manifold is generally twisted and curved; in addition points on it will be non-uniformly distributed: two factors that make the identification of the ID and its exploitation really hard. Here we propose a new ID estimator using only the distance of the first and the second nearest neighbor of each point in the sample. This extreme minimality enables us to reduce the effects of curvature, of density variation, and the resulting computational cost. The ID estimator is theoretically exact in uniformly distributed datasets, and provides consistent measures in general. When used in combination with block analysis, it allows discriminating the relevant dimensions as a function of the block size. This allows estimating the ID even when the data lie on a manifold perturbed by a high-dimensional noise, a situation often encountered in real world data sets. We demonstrate the usefulness of the approach on molecular simulations and image analysis.

Estimating the intrinsic dimension of datasets by a minimal neighborhood information / Facco, Elena; D'Errico, Maria; Rodriguez, Alex; Laio, Alessandro. - In: SCIENTIFIC REPORTS. - ISSN 2045-2322. - 7:1(2017), pp. 1-8. [10.1038/s41598-017-11873-y]

Estimating the intrinsic dimension of datasets by a minimal neighborhood information

Facco, Elena;D'Errico, Maria;Rodriguez, Alex;Laio, Alessandro

2017-01-01

Abstract

Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a manifold whose Intrinsic Dimension (ID) is much lower than the crude large number of coordinates. Such manifold is generally twisted and curved; in addition points on it will be non-uniformly distributed: two factors that make the identification of the ID and its exploitation really hard. Here we propose a new ID estimator using only the distance of the first and the second nearest neighbor of each point in the sample. This extreme minimality enables us to reduce the effects of curvature, of density variation, and the resulting computational cost. The ID estimator is theoretically exact in uniformly distributed datasets, and provides consistent measures in general. When used in combination with block analysis, it allows discriminating the relevant dimensions as a function of the block size. This allows estimating the ID even when the data lie on a manifold perturbed by a high-dimensional noise, a situation often encountered in real world data sets. We demonstrate the usefulness of the approach on molecular simulations and image analysis.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
			2017
		
	Rivista
	
			SCIENTIFIC REPORTS
		
	Numero del volume
	
			7
		
	Fascicolo
	
			1
		
	Da pagina
	
			1
		
	A pagina
	
			8
		
	Numero di articolo
	
			12140
		
	Codice DOI
	
			https://dx.doi.org/10.1038/s41598-017-11873-y
		
	Fulltext via DOI
	
			10.1038/s41598-017-11873-y
		
	URL
	
			https://www.nature.com/articles/s41598-017-11873-y
		
	Tutti gli autori
	
			Facco, Elena; D'Errico, Maria; Rodriguez, Alex; Laio, Alessandro
		
	Appare nelle tipologie:
	
			1.1 Journal article

File in questo prodotto:

File	Dimensione	Formato
s41598-017-11873-y.pdf accesso aperto Tipologia: Versione Editoriale (PDF) Licenza: Creative commons Dimensione 2.62 MB Formato Adobe PDF Visualizza/Apri	2.62 MB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11767/67649

Citazioni

30

144

131

social impact