This paper aims at contributing to the recent debate on the ‘health state’ of Italian language, with a specific focus on the Italian written by Italian university students. The topic has been addressed by the recently closed Univers-ITA project (funded in the PRIN 2017 call) from which the data analysed in this paper come. One of the research lines involved asking a sample of university students, from different Italian universities and different learning subjects, to write a formal text composed of at most 500 words. The texts have been annotated by experts and the number of occurrences of specific features related to orthography, lexicon, syntax, coherence and others have been recorded together with socio-demographic characteristics of the respondents. In order to identify groups of writers sharing the same writing features and to simultaneously allow for the correlation between the features themselves, we propose a new mixture model that models the counts as Poisson random variables whose parameters are generated according to a common factor latent variable model. Identifiability conditions are discussed. The model is applied to the project data.
Dallari, S., Grandi, N., Montanari, A. (2025). What are the Main Features of the Italian Written by Italian University Students?. Cham : Springer [10.1007/978-3-031-64350-7_23].
What are the Main Features of the Italian Written by Italian University Students?
Dallari, Silvia
;Grandi, Nicola;Montanari, Angela
2025
Abstract
This paper aims at contributing to the recent debate on the ‘health state’ of Italian language, with a specific focus on the Italian written by Italian university students. The topic has been addressed by the recently closed Univers-ITA project (funded in the PRIN 2017 call) from which the data analysed in this paper come. One of the research lines involved asking a sample of university students, from different Italian universities and different learning subjects, to write a formal text composed of at most 500 words. The texts have been annotated by experts and the number of occurrences of specific features related to orthography, lexicon, syntax, coherence and others have been recorded together with socio-demographic characteristics of the respondents. In order to identify groups of writers sharing the same writing features and to simultaneously allow for the correlation between the features themselves, we propose a new mixture model that models the counts as Poisson random variables whose parameters are generated according to a common factor latent variable model. Identifiability conditions are discussed. The model is applied to the project data.| File | Dimensione | Formato | |
|---|---|---|---|
|
Paper_SIS_2024.pdf
Open Access dal 04/03/2026
Tipo:
Postprint / Author's Accepted Manuscript (AAM) - versione accettata per la pubblicazione dopo la peer-review
Licenza:
Licenza per accesso libero gratuito
Dimensione
242.58 kB
Formato
Adobe PDF
|
242.58 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



