This study highlights the limitations of common data transformations used to fit transcriptomic data, particularly single-cell RNAseq data, into a normal distribution. Through simulation, it demonstrates that such transformations can distort graphical structures and affect the interpretation of results in two-sample problems, emphasizing the need for specialized methods tailored for count data. - The rapid expansion of technology in medical and biological domains has led to a surge in available data, accompanied by increased complexity. Classical statistical methods, primarily developed for analyzing data assumed to follow a normal distribution, may prove inadequate for handling this complexity. We focus our attention on transcriptomic data, in particular single-cell RNAseq data, which measure gene expression as counts rather than intensities. Through a simulation study, this study aims to demonstrate the limitations of common data transformations intended to fit data into a normal distribution, highlighting potential misinterpretations. Simulation results show that the transformation of the data might affect the results and their interpretation. In particular, we show that in a two-sample problem in a graphical framework, the set of nodes that differs in the two conditions is affected by the transformations. This behavior might be due to the fact that transformations can distort graphical structures reliant on conditional dependencies among variables, affecting final conclusions. Specifically, when analyzing RNA-seq data, methods tailored for count data are preferable over those designed for normally distributed data. The findings underscore the need for specialized approaches in handling count data in two-sample problems and advocate for further research into alternative methods for differential network analysis.

Banzato, E., Risso, D., Chiogna, M., Djordjilović, V. (2025). Data Transformation and Its Validity in a Two-Sample Problem: An Illustration Based on Graphical Models. Springer Nature [10.1007/978-3-031-64431-3_21].

Data Transformation and Its Validity in a Two-Sample Problem: An Illustration Based on Graphical Models

Erika Banzato
;
Monica Chiogna;
2025

Abstract

This study highlights the limitations of common data transformations used to fit transcriptomic data, particularly single-cell RNAseq data, into a normal distribution. Through simulation, it demonstrates that such transformations can distort graphical structures and affect the interpretation of results in two-sample problems, emphasizing the need for specialized methods tailored for count data. - The rapid expansion of technology in medical and biological domains has led to a surge in available data, accompanied by increased complexity. Classical statistical methods, primarily developed for analyzing data assumed to follow a normal distribution, may prove inadequate for handling this complexity. We focus our attention on transcriptomic data, in particular single-cell RNAseq data, which measure gene expression as counts rather than intensities. Through a simulation study, this study aims to demonstrate the limitations of common data transformations intended to fit data into a normal distribution, highlighting potential misinterpretations. Simulation results show that the transformation of the data might affect the results and their interpretation. In particular, we show that in a two-sample problem in a graphical framework, the set of nodes that differs in the two conditions is affected by the transformations. This behavior might be due to the fact that transformations can distort graphical structures reliant on conditional dependencies among variables, affecting final conclusions. Specifically, when analyzing RNA-seq data, methods tailored for count data are preferable over those designed for normally distributed data. The findings underscore the need for specialized approaches in handling count data in two-sample problems and advocate for further research into alternative methods for differential network analysis.
2025
Methodological and Applied Statistics and Demography III
121
126
Banzato, E., Risso, D., Chiogna, M., Djordjilović, V. (2025). Data Transformation and Its Validity in a Two-Sample Problem: An Illustration Based on Graphical Models. Springer Nature [10.1007/978-3-031-64431-3_21].
Banzato, Erika; Risso, Davide; Chiogna, Monica; Djordjilović, Vera
File in questo prodotto:
File Dimensione Formato  
sis_2024.pdf

Open Access dal 31/01/2026

Tipo: Postprint / Author's Accepted Manuscript (AAM) - versione accettata per la pubblicazione dopo la peer-review
Licenza: Licenza per accesso libero gratuito
Dimensione 265.74 kB
Formato Adobe PDF
265.74 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11585/1015426
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? 0
  • OpenAlex ND
social impact