CRIS Current Research Information System

Text-to-audio synthesis (TTA) promises sound generation mediated solely by natural language. Like all generative deep learning techniques, it derives its world memory and semantic boundaries from the dataset with which it is trained. Following a humanistic tradition of critical algorithm analysis, this article examines the datasets employed by various text-to-audio synthesis models and how they are designed and manipulated by researchers. Datasets are not inert informational structures, but socio-technical objects involving specific symbolisations mediated by cultural, po litical, and industrial perspectives. A comparative study of label datasets, caption datasets, and algorithmically augmented datasets reveals the technical and ethical limitations associated with a quantitative approach to data collection. In contrast to a purely computational dataset evaluation paradigm, a qualitative analysis methodology is proposed to interpret the connotative capability of models and suggest alternative practices in the collection of information.

Ancona, R. (2024). Una prospettiva critica sui dataset per la sintesi text-to-audio.

Una prospettiva critica sui dataset per la sintesi text-to-audio

Riccardo Ancona

2024

Abstract

Text-to-audio synthesis (TTA) promises sound generation mediated solely by natural language. Like all generative deep learning techniques, it derives its world memory and semantic boundaries from the dataset with which it is trained. Following a humanistic tradition of critical algorithm analysis, this article examines the datasets employed by various text-to-audio synthesis models and how they are designed and manipulated by researchers. Datasets are not inert informational structures, but socio-technical objects involving specific symbolisations mediated by cultural, po litical, and industrial perspectives. A comparative study of label datasets, caption datasets, and algorithmically augmented datasets reveals the technical and ethical limitations associated with a quantitative approach to data collection. In contrast to a purely computational dataset evaluation paradigm, a qualitative analysis methodology is proposed to interpret the connotative capability of models and suggest alternative practices in the collection of information.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2024
			
	Titolo del volume
	
				Memorie proiettive/Projecting Memories. XXIV CIM – Colloquio di Informatica Musicale
			
	Pagina iniziale
	
				107
			
	Pagina finale
	
				114
			
	Citazione
	
				Ancona, R. (2024). Una prospettiva critica sui dataset per la sintesi text-to-audio.
			
	Tutti gli autori
	
						Ancona, Riccardo
					
	Appare nelle tipologie:
	
				4.01 Contributo in Atti di convegno

File in questo prodotto:

File	Dimensione	Formato
ANCONA_2024_CIM_XXIV_Atti (1).pdf accesso aperto Tipo: Versione (PDF) editoriale / Version Of Record Licenza: Licenza per Accesso Aperto. Creative Commons Attribuzione - Non commerciale - Non opere derivate (CCBYNCND) Dimensione 5.02 MB Formato Adobe PDF Visualizza/Apri	5.02 MB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11585/1010833

Citazioni

ND

ND

ND

ND

social impact