Wikipedia’s content is based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close this gap, we release Wikipedia Citations, a comprehensive data set of citations extracted from Wikipedia. We extracted 29.3 million citations from 6.1 million English Wikipedia articles as of May 2020, and classified as being books, journal articles, or Web content. We were thus able to extract 4.0 million citations to scholarly publications with known identifiers—including DOI, PMC, PMID, and ISBN—and further equip an extra 261 thousand citations with DOIs from Crossref. As a result, we find that 6.7% of Wikipedia articles cite at least one journal article with an associated DOI, and that Wikipedia cites just 2% of all articles with a DOI currently indexed in the Web of Science. We release our code to allow the community to extend upon our work and update the data set in the future.

Wikipedia citations: A comprehensive data set of citations with identifiers extracted from english wikipedia / Singh Harshdeep; West Robert; Colavizza Giovanni. - In: QUANTITATIVE SCIENCE STUDIES. - ISSN 2641-3337. - ELETTRONICO. - 2:1(2021), pp. 1-19. [10.1162/qss_a_00105]

Wikipedia citations: A comprehensive data set of citations with identifiers extracted from english wikipedia

Colavizza Giovanni
2021

Abstract

Wikipedia’s content is based on reliable and published sources. To this date, relatively little is known about what sources Wikipedia relies on, in part because extracting citations and identifying cited sources is challenging. To close this gap, we release Wikipedia Citations, a comprehensive data set of citations extracted from Wikipedia. We extracted 29.3 million citations from 6.1 million English Wikipedia articles as of May 2020, and classified as being books, journal articles, or Web content. We were thus able to extract 4.0 million citations to scholarly publications with known identifiers—including DOI, PMC, PMID, and ISBN—and further equip an extra 261 thousand citations with DOIs from Crossref. As a result, we find that 6.7% of Wikipedia articles cite at least one journal article with an associated DOI, and that Wikipedia cites just 2% of all articles with a DOI currently indexed in the Web of Science. We release our code to allow the community to extend upon our work and update the data set in the future.
2021
Wikipedia citations: A comprehensive data set of citations with identifiers extracted from english wikipedia / Singh Harshdeep; West Robert; Colavizza Giovanni. - In: QUANTITATIVE SCIENCE STUDIES. - ISSN 2641-3337. - ELETTRONICO. - 2:1(2021), pp. 1-19. [10.1162/qss_a_00105]
Singh Harshdeep; West Robert; Colavizza Giovanni
File in questo prodotto:
File Dimensione Formato  
Singh et al. - 2021 - Wikipedia citations A comprehensive data set of c.pdf

accesso aperto

Descrizione: Articolo
Tipo: Versione (PDF) editoriale
Licenza: Licenza per Accesso Aperto. Creative Commons Attribuzione (CCBY)
Dimensione 760.96 kB
Formato Adobe PDF
760.96 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11585/948750
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 16
  • ???jsp.display-item.citation.isi??? 12
social impact