A Foundational Protocol for Reproducible Visualization in Multivariate Quantum Data ()
1. Introduction
The curse of dimensionality presents a fundamental barrier to understanding complex systems [1]. In fields such as genomics and quantum many-body physics [2] [3], the state of a system is described by hundreds or thousands of interdependent variables, creating a data landscape that is intrinsically high-dimensional and difficult to comprehend in its raw form. Visualization—the projection of data into a 2D or 3D space accessible to human comprehension—is therefore not merely an aid but a necessity for scientific discovery. Dimensionality reduction (DR) provides the mathematical foundation for this process, serving as the essential bridge between abstract, high-dimensional data and actionable insight [4].
While linear DR methods like Principal Component Analysis (PCA) [5] are computationally robust and provide a unique, reproducible projection, they face a critical limitation: they assume the data’s intrinsic structure lies in a linear subspace. This assumption fails for the intricate, non-linear interactions that define many modern scientific domains. The coordinated expression of gene networks [6] [7], the entangled wavefunctions of quantum particles [8], or the emergent phenomena in complex materials [9] all manifest as complex, low-dimensional manifolds embedded within a high-dimensional space. To faithfully visualize these structures, non-linear dimensionality reduction (NLDR) is not just beneficial—it is crucial. Techniques like UMAP (Uniform Manifold Approximation and Projection) are specifically designed to unravel and preserve these non-linear relationships, making the invisible fabric of complex interactions visually apparent [10]-[13]. Other widely used NLDR methods, such as t-SNE, pursue similar goals of manifold visualization but are not the focus of the present work.
However, this powerful capability introduces a profound methodological challenge. Unlike their linear counterparts, NLDR methods are inherently stochastic and sensitive to hyperparameters. A UMAP embedding is the result of a non-convex optimization process, meaning that different runs on the same data can converge to different local minima, producing visually distinct—and sometimes contradictory—projections. This lack of determinism creates a crisis of reproducibility: if a visualization cannot be consistently reproduced, can it be trusted as a foundation for scientific reasoning? When one research group identifies a “cluster” or a “trajectory” in their UMAP plot, while another sees a different organization, there is no objective framework to determine which, if either, reflects a true property of the data versus an artifact of the algorithm’s randomness [14] [15].
This paper addresses this critical gap by establishing a rigorous, protocol-driven framework for reproducible NLDR. We posit that for NLDR to be a reliable tool for scientific visualization, it must transcend its role as an exploratory heuristic and adopt the reproducibility standards expected of any scientific measurement. We introduce a mathematically defined UMAP protocol that establishes formal convergence criteria, transforming subjective visual assessment into objective, quantitative validation. Through a systematic study of high-dimensional quantum data, we demonstrate how preprocessing and parameter selection dictate the pathway to reproducibility, providing clear guidelines for ensuring that visual discoveries are not stochastic artifacts but robust, verifiable features of the underlying data. Our work aims to fortify the crucial bridge between data and discovery, ensuring that the power of non-linear visualization is matched by the rigor required for conclusive scientific interpretation.
2. Data Organization and Structure
This work leverages high-dimensional quantum data to establish a rigorous visualization protocol, with a direct application to documenting the novel nonergodic metal (NEM) phase [16]. As visual confirmation is crucial for interpreting complex phase transitions, as done in [17], according to the pipeline presented in Figure 1, which produces an unsupervised machine learning phase classifier. Our objective with the current work is to produce reliable, human-interpretable projections of the underlying quantum state space.
The core data for each interaction strength
forms a matrix
, where
rows correspond to unique quantum eigenstates sampled from a trajectory across the generalized Aubry-André model, and
columns are the ordered components of each eigenstate’s entanglement spectrum [18]. Our fundamental challenge is the dimensionality reduction
, where
is a 2D embedding suitable for visual analysis.
Figure 1. Research pipeline for quantum phase classification, with post-visualization.
The feature space exhibits a pronounced multiscale character, spanning an immense dynamic range. This is illustrated in Figure 2, which presents the median trajectory for each feature using all sample space for the whole
datasets, plotted on a logarithmic scale. The highest feature values are on the order of 100, while the lowest values approach 10−17. To contextualize this scale, the difference in variance between the dominant and subtle features is comparable to the difference in mass between a school bus and the planet Pluto.
Figure 2. Feature-wise mean across all
, presented on a logarithmic scale.
3. A Reproducibility Protocol for UMAP
To establish UMAP as a reliable component of the scientific process for analyzing quantum phenomena, it is imperative to guarantee that its visual outputs are not artifacts of stochastic optimization but faithful, reproducible representations of the underlying data structure. We introduce a formal protocol centered on a rigorous definition of embedding convergence, which provides the necessary foundation for trustworthy visual analysis.
3.1. UMAP Operation
Uniform Manifold Approximation and Projection (UMAP) [13] constructs a low-dimensional embedding through a two-stage process. First, it builds a fuzzy topological representation of the high-dimensional data by identifying neighborhoods through a k-nearest neighbor (KNN) graph. The parameter
(n_neighbors) controls the scale of this construction: small
values capture fine-grained local structure, while larger
values emphasize broader global relationships.
The second stage employs stochastic gradient descent to optimize a low-dimensional layout that preserves this topological structure. The optimization applies attractive forces between neighboring points and repulsive forces between non-neighbors, minimizing a cross-entropy loss function. This stochastic process, combined with random initialization, means that independent runs can produce visually distinct embeddings from the same data and parameters.
3.2. The Significance of Convergent Embeddings
The core challenge in applying stochastic dimensionality reduction techniques like UMAP to scientific discovery is distinguishing meaningful geometric structure from computational artifacts. A UMAP embedding that varies significantly across runs with identical parameters offers little scientific value, as its visual representation cannot be reliably interpreted or communicated.
The ideal scenario is convergence to a unique and stable geometry. However, this is complicated by large sample sizes and extreme dynamic ranges in multivariate feature spaces, both common characteristics of scientific datasets. When multiple independent executions of UMAP under a fixed configuration nonetheless produce identical connected structures, it provides compelling evidence that these structures are intrinsic properties of the preprocessed data’s topology rather than stochastic artifacts. Achieving this convergent state—where the number of connected substructures is consistently one across all runs with zero variance—transforms UMAP from an exploratory heuristic into a deterministic mapping procedure.
This convergence is the cornerstone of scientific utility. Once a unique embedding is guaranteed for a given dataset and preprocessing method, it can serve as a definitive visual map for consistently interpreting and explaining that dataset. Different scientists can independently arrive at the same visualization, enabling unified interpretation, direct comparison of results, and collaborative hypothesis testing on a stable geometric representation. Furthermore, by achieving this convergent state for both raw and standardized data, one can then visually inspect whether the two preprocessing pathways reveal the same underlying global structure, providing a powerful cross-validation of the observed phenomena.
3.3. Mathematical Framework for Convergence
To formalize this notion, we define a discrete random variable
representing the number of connected substructures in an embedding. For a fixed UMAP parameter set
and preprocessing function
, we execute
independent runs to obtain embeddings
. For each embedding,
is the number of connected components in its KNN graph for a given N.
Our primary theoretical claim is that for a suitably preprocessed dataset, the stochasticity of UMAP can be controlled via the neighborhood parameter
. Specifically, we posit that
converges in probability to 1 as
increases:
(1)
Justification: As
increases, the number of neighborhoods with a population
decreases, and the k-NN graph underlying UMAP’s initial topological representation becomes less fractioned (more connected). In the limiting case where
approaches the sample size, the graph becomes fully connected, possessing exactly one connected component. The optimization process, while stochastic, aims to find a low-dimensional layout that preserves this connectivity structure. For sufficiently large
, the attractive forces imposed by the global connectivity dominate the repulsive forces, constraining the optimization to produce a single, coherent structure. Standardization facilitates this convergence by ensuring the graph’s connectivity reflects balanced global correlations rather than being fragmented by disparate feature scales.
Based on this convergence behavior, we define a rigid, empirical criterion for a configuration to be considered minimally reproducible in practice. A configuration
is necessary reproducible if, across a finite number of runs
, the following condition is met:
(2)
This condition ensures that every execution produces a single, connected structure with zero empirical variance, guaranteeing a unique connected geometry for scientific analysis.
3.4. Protocol Implementation
The implementation of our reproducibility protocol follows a rigorous pipeline designed to systematically evaluate stability across the entire phase space of the quantum model:
1) Systematic Data Generation: The protocol is applied across the full spectrum of quantum phases. We generate datasets for multiple interaction strengths,
, ensuring the analysis covers diverse quantum regimes.
2) Preprocessing: For each
, we generate both the raw dataset
and the standardized dataset
.
3) Multi-Scale Screening: For each dataset
and for each
, we execute
independent UMAP runs. The choice of
represents a balance between computational feasibility and statistical sensitivity: given the binary nature of the convergence criterion (zero versus non-zero variance in the number of connected components), ten trials are sufficient to reliably detect empirical instability while keeping the total computational cost manageable for large datasets.
4) Consistency Assessment: For each configuration
, we compute the empirical distribution of
across the
runs and apply the convergence criteria from Equation (2).
5) Convergence Verification: For each
and preprocessing method
, we identify the minimal
for which the configuration achieves scientific reproducibility.
6) Robustness Validation: We aggregate results across all
values. A preprocessing method
is considered universally robust if it achieves scientific reproducibility for all
at a consistent and feasible
.
This protocol provides a mathematically grounded framework to identify UMAP configurations that converge to a unique, stable embedding. The subsequent application of this protocol determines which configurations, if any, meet the criteria for scientific reproducibility, thereby objectively evaluating UMAP’s suitability for rigorous visual analysis of high-dimensional quantum data.
4. Methodology and Results for Substructure Convergence Analysis
To mirror the default behaviour of exploratory data analysis, each UMAP execution was initialized using the algorithm’s random initialization (i.e., without a fixed seed). As a result, the stochastic gradient descent trajectory differs across runs, making
an empirical random variable that quantifies variability in the embedding’s substructures as a function of
. We estimate the distribution of
from
independent trials.
This section reports the results of applying our reproducibility protocol to nine multivariate quantum datasets spanning interaction strengths
to 0.9, with comprehensive statistics for the number of connected substructures summarized in Table 1. Across all interaction regimes, the number of disconnected substructures—measured as connected components in the two-dimensional KNN graph—decreases monotonically with increasing
, consistent with theoretical expectations that larger neighborhoods enhance global connectivity by expanding UMAP’s local simplicial complexes.
Table 1. Empirical distribution of the random variable
(number of UMAP connected components) comparing raw and standardized data for neighborhood sizes
. Bold italic values indicate fully connected embeddings (Mean = 1, Var. = 0).
Dataset |
N |
Raw Data |
Standard Scaled |
|
|
Min |
Max |
Mean |
Var. |
Min |
Max |
Mean |
Var. |
|
10 |
1 |
3 |
1.9 |
0.49 |
1 |
3 |
1.6 |
0.44 |
15 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
20 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
30 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
40 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
50 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
|
10 |
3 |
5 |
3.9 |
0.49 |
1 |
3 |
2.0 |
0.59 |
15 |
2 |
2 |
2.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
20 |
1 |
2 |
1.4 |
0.24 |
1 |
1 |
1.0 |
0.00 |
30 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
40 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
50 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
|
10 |
12 |
16 |
14.5 |
1.85 |
1 |
2 |
1.4 |
0.24 |
15 |
4 |
7 |
5.7 |
0.81 |
1 |
1 |
1.0 |
0.00 |
20 |
3 |
5 |
4.2 |
0.76 |
1 |
1 |
1.0 |
0.00 |
30 |
2 |
2 |
2.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
40 |
2 |
2 |
2.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
50 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
|
10 |
20 |
32 |
25.9 |
8.29 |
3 |
5 |
3.5 |
0.45 |
15 |
6 |
9 |
7.8 |
1.56 |
1 |
1 |
1.0 |
0.00 |
20 |
2 |
4 |
2.4 |
0.44 |
1 |
2 |
1.1 |
0.09 |
30 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
40 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
50 |
1 |
1 |
1.0 |
0.00 |
1 |
1 |
1.0 |
0.00 |
|
10 |
31 |
40 |
36.1 |
5.48 |
3 |
7 |
4.5 |
1.44 |
15 |
15 |
20 |
17.8 |
1.96 |
1 |
2 |
1.1 |
0.09 |
20 |
10 |
17 |
13.0 |
2.99 |
1 |
1 |
1.0 |
0.00 |
30 |
3 |
7 |
5.0 |
1.21 |
1 |
1 |
1.0 |
0.00 |
40 |
2 |
4 |
3.0 |
0.40 |
1 |
1 |
1.0 |
0.00 |
50 |
3 |
4 |
3.4 |
0.24 |
1 |
1 |
1.0 |
0.00 |
|
10 |
37 |
49 |
42.8 |
13.54 |
4 |
8 |
6.0 |
1.21 |
15 |
10 |
23 |
18.2 |
15.76 |
2 |
3 |
2.2 |
0.16 |
20 |
5 |
10 |
7.3 |
1.82 |
2 |
2 |
2.0 |
0.00 |
30 |
2 |
4 |
3.3 |
0.61 |
1 |
2 |
1.9 |
0.09 |
40 |
2 |
4 |
3.1 |
0.69 |
1 |
2 |
1.1 |
0.09 |
50 |
2 |
3 |
2.3 |
0.21 |
1 |
1 |
1.0 |
0.00 |
|
10 |
69 |
83 |
75.5 |
16.24 |
13 |
20 |
15.5 |
3.84 |
15 |
34 |
44 |
40.2 |
9.55 |
1 |
3 |
2.3 |
0.41 |
20 |
20 |
23 |
22.0 |
1.00 |
1 |
4 |
2.1 |
0.69 |
30 |
10 |
14 |
12.0 |
1.80 |
1 |
2 |
1.3 |
0.21 |
40 |
3 |
11 |
6.9 |
5.29 |
1 |
1 |
1.0 |
0.00 |
50 |
3 |
4 |
3.1 |
0.09 |
1 |
1 |
1.0 |
0.00 |
|
10 |
71 |
84 |
77.9 |
22.85 |
15 |
20 |
17.4 |
3.24 |
15 |
32 |
44 |
36.5 |
11.83 |
2 |
6 |
4.0 |
1.21 |
20 |
15 |
23 |
18.2 |
5.38 |
1 |
2 |
1.2 |
0.16 |
30 |
9 |
13 |
11.1 |
1.69 |
1 |
1 |
1.0 |
0.00 |
40 |
1 |
5 |
3.3 |
1.00 |
1 |
1 |
1.0 |
0.00 |
50 |
1 |
3 |
1.5 |
0.45 |
1 |
1 |
1.0 |
0.00 |
|
10 |
82 |
93 |
87.8 |
11.56 |
14 |
22 |
19.6 |
4.45 |
15 |
27 |
40 |
34.2 |
11.16 |
2 |
7 |
4.6 |
1.85 |
20 |
16 |
27 |
22.6 |
10.24 |
1 |
2 |
1.2 |
0.16 |
30 |
6 |
11 |
7.8 |
3.17 |
1 |
1 |
1.0 |
0.00 |
40 |
4 |
7 |
5.6 |
0.64 |
1 |
1 |
1.0 |
0.00 |
50 |
2 |
4 |
2.6 |
0.44 |
1 |
1 |
1.0 |
0.00 |
For raw data, small neighborhood sizes (
) produce highly fragmented embeddings, often with tens of disconnected components at intermediate interaction strengths (
). These regimes correspond to unstable manifolds that are insufficiently sampled, where stochastic initialization yields qualitatively distinct topologies across runs, reflected in high variance. As
increases, connectivity constraints are relaxed and components merge, reducing stochastic sensitivity and stabilizing the embedding.
These results indicate that substructure convergence in UMAP can be interpreted as a discrete topological transition governed by the connectivity of the neighborhood graph: fragmented simplicial complexes at small
progressively merge until a single connected manifold emerges. The zero-variance condition marks the onset of global topological stability under random initialization and parameter noise.
In summary, convergence toward a single connected embedding is a general property of UMAP on complex multivariate datasets. This process is accelerated and stabilized by standardization, which enforces uniform metric scaling, and is guaranteed at sufficiently large neighborhood sizes. The resulting criterion provides a principled method for identifying the minimal
required for topologically stable and reproducible embeddings, offering a rigorous foundation for UMAP-based analysis in mathematical and quantum physics contexts.
4.1. Standardized Data Produces a Single Connected Structure
For standardized data, UMAP embeddings consistently converge to a single connected structure across all nine quantum datasets when the neighborhood size
is sufficiently large. In most datasets (
), this convergence occurs already at
, while for all datasets
guarantees a single structure. The variance column for standardized data in Table 1 shows a clear trend: as
increases, the number of disconnected substructures systematically reduces, ultimately reaching one. This behavior confirms that standardization stabilizes local and global relationships in high-dimensional quantum data, enabling reproducible structural interpretations.
4.2. Fragmentation in Raw Data Embeddings
In contrast, embeddings generated from raw data exhibit pronounced fragmentation and are highly sensitive to initialization. At small neighborhood sizes (
), the mean number of disconnected substructures ranges from 1.9 (
) to 87.8 (
), with substantial variance, indicating highly irregular and fractured embeddings. Increasing
progressively reduces fragmentation, but a single connected structure is achieved only in 4 of 9 datasets
(
). For intermediate regimes (
) and higher interaction strengths (
), multiple disconnected substructures persist even at
, demonstrating that raw data cannot reliably produce a fully connected embedding in all cases.
4.3. Preliminary Insights for Reliable Non-Linear Dimensionality Reduction
The pursuit of reliable visualization for complex multivariate quantum data presents a significant challenge. This investigation, conducted across nine large-scale datasets (approximately 23,000 samples with 750 features), explored the stability of UMAP embeddings as a potential path forward. The experimental protocol, involving 10 independent executions for each neighborhood parameter (N), provides initial evidence, though further validation is needed.
A central finding from these preliminary tests, summarized in Table 2, concerns the optimization landscape. The raw data’s high-dimensional, fragmented nature appears to create a complex topography with many local minima, slowing convergence, and often leading to multiple substructures. In contrast, standardized data facilitates a more direct path to a single, deterministic embedding by presenting a smoother global structure for the algorithm to capture.
Table 2. Convergence results across datasets.
|
|
Std. |
|
Raw |
0.1 |
15 |
√ |
15 |
√ |
0.2 |
20 |
√ |
– |
× |
0.3 |
30 |
√ |
– |
× |
0.4 |
30 |
√ |
– |
× |
0.5 |
30 |
√ |
– |
× |
0.6 |
40 |
√ |
50 |
√ |
0.7 |
40 |
√ |
50 |
√ |
0.8 |
40 |
√ |
50 |
√ |
0.9 |
50 |
√ |
50 |
√ |
Total |
– |
9/9 |
– |
4/9 |
These observations begin to outline a methodological approach for generating more reliable explanations of quantum data through visualization:
Data Preprocessing as a Stabilizer: The consistent performance of standardized data suggests that preprocessing which promotes global coherence can be a powerful tool for stabilizing the embedding process and mitigating local fragmentation.
Navigating the Optimization Landscape: The results indicate that the choice of neighborhood size is pivotal. Larger neighborhood sizes appear to help the algorithm navigate past local minima in complex data, a consideration that becomes especially relevant when preprocessing is not applied.
Convergence as a Diagnostic Metric: Monitoring variance in substructure counts across multiple runs is a practical diagnostic. A consistent, single structure across executions can bolster confidence that the resulting visualization reflects a stable global geometry rather than a transient local configuration.
Interpreting with an Awareness of Stability: Embeddings that are sensitive to initial conditions may offer valuable but partial insights. Acknowledging this instability can lead to a more nuanced interpretation, where multiple runs are seen as exploring a landscape of possible structures inherent to the data.
This preliminary work suggests that a conscious approach to data preparation and parameter selection, coupled with simple validation checks, can pave the way for more trustworthy visualizations. For complex multivariate quantum data, such practices may ultimately lead to more reliable explanations and a deeper understanding of the underlying quantum phenomena.
4.4. Limitations and Future Work
While our protocol establishes a necessary condition for reproducible visualization—the convergence to a stable, connected embedding—we acknowledge that reproducibility alone does not guarantee physical interpretability. A legitimate critique is that by increasing the neighborhood parameter
or applying standardization, we might be artificially connecting genuinely distinct physical regions, effectively smoothing over physically significant discontinuities.
However, we posit that achieving a single, reproducible structure is a fundamental prerequisite for any robust physical interpretation. The high fragmentation observed in raw data embeddings at small
values, characterized by high variance across runs, represents a state of analytical chaos where no stable interpretation is even possible. Our protocol systematically navigates this stochasticity to arrive at a deterministic endpoint. The resulting connected structure provides a stable canvas upon which physically meaningful patterns (such as gradients, clusters, or topological voids) can be reliably identified and studied.
In this work, we focus on establishing this methodological foundation of reproducibility. A comprehensive validation of the physical meaning embedded within these stable structures—a visual cross-referencing of UMAP patterns with independent physical observables—remains a critical task for future work. Such an analysis is challenging in this context due to the lack of a clear ground truth and the high dimensionality of the original space, which complicates quantitative measures such as trustworthiness and continuity. Future research will involve a detailed visual comparison across the quantum phase diagram and the development of specialized metrics to quantify how well the embedding preserves the physical relationships between a representative subset of states. Future work will incorporate pseudo ground-truth physical observables as color overlays on the converged embeddings, enabling a direct visual cross-validation and providing a more suitable and intuitive tool for physical inspection.
5. Conclusions
This work establishes a rigorous protocol for rendering UMAP a reproducible tool in scientific visualization. By framing reproducibility as the convergence toward a single, connected embedding topology, we introduced quantitative criteria that transform a stochastic algorithm into a verifiable mapping process. Applied to complex multivariate quantum datasets, our protocol revealed a consistent and interpretable pattern: as the neighborhood parameter
increases, the embedding undergoes a discrete topological transition from fragmented to unified manifolds, converging to a deterministic structure once global connectivity percolates through the data graph.
Standardization plays a decisive role in this transition. By rescaling the features, it equalizes the contribution of disparate energy scales and promotes faster convergence at smaller
, effectively stabilizing the manifold reconstruction. In contrast, raw data require larger neighborhoods to achieve the same level of reproducibility, underscoring the sensitivity of distance metrics to feature-scale imbalance. Across all interaction regimes, however, the limiting behavior is universal: UMAP asymptotically converges to a single, stable embedding once sufficient neighborhood connectivity is reached.
The implications extend beyond the present application. The proposed reproducibility protocol provides a general framework for distinguishing robust visual patterns from stochastic artifacts across domains using non-linear embeddings. It offers a methodological foundation for integrating visualization within the standards of scientific reproducibility, where a figure represents not a random snapshot of computation, but a stable property of the underlying data manifold. Future work will expand this framework to include geometric stability metrics and explore how convergence correlates with physically meaningful invariants in quantum phenomena.