Scientific Maps Analysis with the R package bibliometrix

Example of scientific maps for bibliometric analysis
bibliometry
mapping
bibliometric analysis
Author
Published

Friday, September 9, 2022

Doi

PreregisteredOpen dataOpen

Introduction

Bibliographic Data

The bibliographic data are based on the search query that has been made. For the moment we have focused on one of the lines of the SEJ670 UCO Research Group, neophilia/neophobia in tourism. Needles to say that it would be applied to other research lines; however, watch out when filtering and redefining the database because we have to remove areas that are not of our interest, for example, in our selection there were references about zoology, biology, etc…, given that there is also neophilia and neophobia, especially the latter, specifically in animals, well, we must be careful because some area can sneak in the form of journals and references not related to our research area.

Data source: Clarivate Analytics Web of Science (http://apps.webofknowledge.com)

Format: Bibtex

Query: “Web of Science Core Collection”

Range: 1995-2022

Document type: All

Query date: September, 2022

Package installation, data loading and conversion

As the objective is to see what analysis the tool does and how it does it, the steps of installation, data loading and conversion have been omitted, although it is an important step since it must be done in R, after downloading the data from WoS or Scopus, which also has its own technique. After a lot of research we managed to find a way to merge both databases and it works perfectly, and it also automatically omits duplicate references. With this procedure we can have both databases in a single file and although WoS has more references (as a general rule, although not always), Scopus almost always has some different one/s and even if they are few, at least we practically cover the two most important sources of articles. It is worth mentioning that in the last updates of this Bibliometrix package other open bibliographic databases such as OpenAlex, Dimensions, Lens, and from the medical field such as PubMed and Cochrane have been included.

1. Descriptive Analysis

The descriptive analysis of the package provides a lot of information on the annual development of research, the most productive \(k\) authors, articles, countries and relevant keywords.

1.1. Main findings on the database analysed

The following table shows a summary of the data and other interesting classifications: n_umber of authors, number of documents, scientific production for each year (and its average growth), the most productive authors, the countries and the corresponding citations, the most relevant journals, the most relevant keywords, etc…

Moreover, this information can also be obtained in graphs.



MAIN INFORMATION ABOUT DATA

 Timespan                              1995 : 2022 
 Sources (Journals, Books, etc)        69 
 Documents                             152 
 Annual Growth Rate %                  9.64 
 Document Average Age                  9.52 
 Average citations per doc             29.24 
 Average citations per year per doc    2.425 
 References                            7867 
 
DOCUMENT TYPES                     
 article                         130 
 article; book chapter           2 
 article; early access           5 
 article; proceedings paper      1 
 editorial material              4 
 letter                          1 
 proceedings paper               3 
 review                          6 
 
DOCUMENT CONTENTS
 Keywords Plus (ID)                    617 
 Author's Keywords (DE)                512 
 
AUTHORS
 Authors                               467 
 Author Appearances                    531 
 Authors of single-authored docs       18 
 
AUTHORS COLLABORATION
 Single-authored docs                  19 
 Documents per Author                  0.325 
 Co-Authors per Doc                    3.49 
 International co-authorships %        36.18 
 

Annual Scientific Production

 Year    Articles
    1995        1
    2000        2
    2003        2
    2005        3
    2006        2
    2007        3
    2008        5
    2009        3
    2010        7
    2011        2
    2012        6
    2013        5
    2014        8
    2015       11
    2016       11
    2017       13
    2018       13
    2019       11
    2020       17
    2021       10
    2022       12

Annual Percentage Growth Rate 9.64 


Most Productive Authors

     Authors        Articles   Authors        Articles Fractionalized
1  METTKE-HOFMANN C        8 METTKE-HOFMANN C                    3.53
2  BUGNYAR T               4 STHAPIT E                           1.33
3  EVES A                  4 EVES A                              1.20
4  MORAND-FERRON J         4 MORAND-FERRON J                     1.20
5  BISAZZA A               3 [ANONYMOUS] A                       1.00
6  CARACCIOLO F            3 AKYUZ BG                            1.00
7  GRIFFIN AS              3 ANTONAKIS J                         1.00
8  KIM YG                  3 ARMELAGOS GJ                        1.00
9  VERNEAU F               3 BERTI I                             1.00
10 WIDDIG A                3 CHAKRABARTI B                       1.00


Top manuscripts per citations

                                 Paper                                      DOI  TC TCperYear  NTC
1  KIM YG, 2009, INT J HOSP MANAG               10.1016/j.ijhm.2008.11.005      346     20.35 2.38
2  CHANG RCY, 2010, ANN TOURIS RES              10.1016/j.annals.2010.03.007    276     17.25 3.21
3  ANTONAKIS J, 2017, LEADERSH Q                10.1016/j.leaqua.2017.01.006    192     21.33 4.79
4  VERBEKE W, 2005, BR FOOD J                   10.1108/00070700510629779       183      8.71 1.97
5  KIM YG, 2010, INT J HOSP MANAG               10.1016/j.ijhm.2009.10.015      178     11.12 2.07
6  CAVIGELLI SA, 2003, PROC NATL ACAD SCI U S A 10.1073/pnas.2535721100         161      7.00 1.22
7  HUGHES RN, 2007, NEUROSCI BIOBEHAV REV       10.1016/j.neubiorev.2006.11.004 110      5.79 2.84
8  MORAND-FERRON J, 2011, BEHAV ECOL            10.1093/beheco/arr120           106      7.07 1.49
9  DAY RL, 2003, ANIM BEHAV                     10.1006/anbe.2003.2074          103      4.48 0.78
10 REVERDY C, 2008, APPETITE                    10.1016/j.appet.2008.01.010     101      5.61 1.51


Corresponding Author's Countries

          Country Articles   Freq SCP MCP MCP_Ratio
1  UNITED KINGDOM       24 0.1622  17   7     0.292
2  USA                  18 0.1216  12   6     0.333
3  CHINA                14 0.0946   5   9     0.643
4  AUSTRALIA            13 0.0878   9   4     0.308
5  GERMANY               9 0.0608   4   5     0.556
6  CANADA                8 0.0541   5   3     0.375
7  ITALY                 7 0.0473   7   0     0.000
8  AUSTRIA               6 0.0405   2   4     0.667
9  BELGIUM               3 0.0203   1   2     0.667
10 BRAZIL                3 0.0203   2   1     0.333


SCP: Single Country Publications

MCP: Multiple Country Publications


Total Citations per Country

      Country      Total Citations Average Article Citations
1  UNITED KINGDOM             1122                     46.75
2  USA                         639                     35.50
3  CHINA                       490                     35.00
4  AUSTRALIA                   275                     21.15
5  BELGIUM                     193                     64.33
6  SWITZERLAND                 192                     96.00
7  AUSTRIA                     187                     31.17
8  GERMANY                     158                     17.56
9  ARGENTINA                   143                     71.50
10 ITALY                       143                     20.43


Most Relevant Sources

                                    Sources        Articles
1  ANIMAL BEHAVIOUR                                      17
2  ETHOLOGY                                               8
3  ANIMAL COGNITION                                       7
4  FOOD QUALITY AND PREFERENCE                            7
5  BEHAVIORAL ECOLOGY                                     6
6  BRITISH FOOD JOURNAL                                   6
7  SCIENTIFIC REPORTS                                     6
8  APPETITE                                               5
9  INTERNATIONAL JOURNAL OF HOSPITALITY MANAGEMENT        5
10 PLOS ONE                                               5


Most Relevant Keywords

   Author Keywords (DE)      Articles Keywords-Plus (ID)     Articles
1             NEOPHILIA            24           NEOPHOBIA          48
2             FOOD NEOPHOBIA       23           BEHAVIOR           29
3             NEOPHOBIA            22           NEOPHILIA          21
4             EXPLORATION          13           ATTITUDES          17
5             FOOD                 11           SCALE              17
6             FOOD TOURISM         11           SATISFACTION       16
7             FOOD NEOPHILIA       10           TOURISM            15
8             PERSONALITY           9           EVOLUTION          13
9             INNOVATION            8           INFORMATION        13
10            LOCAL FOOD            8           MODEL              13

1.2. Most cited references

We can also obtain the most cited references and their number:

                                                                                   [,1]
PLINER P, 1992, APPETITE, V19, P105, DOI 10.1016/0195-6663(92)90014-W                55
GREENBERG R, 2001, CURR ORNITHOL, V16, P119                                          32
COHEN E, 2004, ANN TOURISM RES, V31, P755, DOI 10.1016/J.ANNALS.2004.02.003          29
RITCHEY PN, 2003, APPETITE, V40, P163, DOI 10.1016/S0195-6663(02)00134-4             27
METTKE-HOFMANN C, 2002, ETHOLOGY, V108, P249, DOI 10.1046/J.1439-0310.2002.00773.X   24
TUORILA H, 2001, FOOD QUAL PREFER, V12, P29, DOI 10.1016/S0950-3293(00)00025-2       24
KIM YG, 2009, INT J HOSP MANAG, V28, P423, DOI 10.1016/J.IJHM.2008.11.005            21
KIM YG, 2010, INT J HOSP MANAG, V29, P216, DOI 10.1016/J.IJHM.2009.10.015            19
QUAN S, 2004, TOURISM MANAGE, V25, P297, DOI 10.1016/S0261-5177(03)00130-4           19
REALE D, 2007, BIOL REV, V82, P291, DOI 10.1111/J.1469-185X.2007.00010.X             19
CHANG RCY, 2010, ANN TOURISM RES, V37, P989, DOI 10.1016/J.ANNALS.2010.03.007        16
PLINER P, 2006, FRONT NUTR SCI, P75, DOI 10.1079/9780851990323.0075                  16
FORNELL C, 1981, J MARKETING RES, V18, P39, DOI 10.2307/3151312                      15
JI MJ, 2016, TOURISM MANAGE, V57, P387, DOI 10.1016/J.TOURMAN.2016.06.003            15
MAK AHN, 2012, INT J HOSP MANAG, V31, P928, DOI 10.1016/J.IJHM.2011.10.012           15
CHANG RCY, 2011, TOURISM MANAGE, V32, P307, DOI 10.1016/J.TOURMAN.2010.02.009        14
FISCHLER C, 1988, SOC SCI INFORM, V27, P275, DOI 10.1177/053901888027002005          14
GREENBERG RUSSELL, 2003, P175                                                        14
MARTIN LB, 2005, BEHAV ECOL, V16, P702, DOI 10.1093/BEHECO/ARI044                    14
SOL D, 2011, PLOS ONE, V6, DOI 10.1371/JOURNAL.PONE.0019535                          14

2: Co-citation analysis structure

Citation analysis is another remarkable tool of the bibliometric analysis offered by the package. It shows the structure of a specific field through the links between nodes (e.g. authors, articles, journal). The interesting option is that the edges can be interpreted differently depending on the type of network, i.e. co-citations, direct citations, bibliographic linking, etc…. This is very useful and can be exploited quite a lot.

Below we have taken the three standard examples shown in the original reference but adapted them to our database.

First, a co-citation network showing the relationships between the cited-referred works (nodes).

Second, a co-citation network that uses the cited journals as the unit of analysis.

The dimensions useful for commenting on co-citation networks are: (i) centrality and peripherality of nodes, (ii) their proximity and distance, (iii) strength of links, (iv) clusters, (v) bridging contributions.

Third, a historiography that is built on direct quotations. It traces the intellectual links in a historical order. The cited works of thousands of authors contained in a collection of published scientific articles are sufficient to reconstruct the historiographical structure of the field, pointing out the basic works in it.

Analysis of co-citations of articles (references)

This is the typical VosViewer graph, in this case for a visualization of the co-cite network. The graph parameters have not been modified to make it clearer but the graph options shown below are mainly for visual fine tuning. Without seeing the code this may be useless but it is interesting to know that it can be tweaked and made more readable. The interesting thing would be to be able to eliminate those co-citations that appear isolated and focus only on those where there are relationships.

  • n = 50 (this function traces the top 50 cited references)

  • type = "fruchterman" (the network layout is generated by the Fruchterman-Reingold algorithm, there is the option of other types of algorithms, although I have not tested them and I do not know if they would work for us)

  • size.cex = TRUE (the size of the vertices is proportional to its degree)

  • size = 20 (maximum size of the vertices)

  • remove.multiple = FALSE (multiple edges are not eliminated, the opposite is TRUE)

  • labelsize = 1 (defines the size of the vertices labels)

  • edgesize = 10 (the thickness of the edges is proportional to their strength. Edgesize defines the maximum value of the thickness)

  • edges.min = 5 (only traces edges with a force greater than or equal to 5)

  • all other arguments assume default values.

Analysis of journal co-citations (source)

wos=metaTagExtraction(wos,"CR_SO",sep=";")
NetMatrix <- biblioNetwork(wos, analysis = "co-citation", network = "sources", sep = ";")
net=networkPlot(NetMatrix, n = 50, Title = "Co-Citation Network", type = "auto", size.cex=TRUE, size=15, remove.multiple=FALSE, labelsize=1,edgesize = 10, edges.min=5)

Descriptive analysis of the characteristics of the journal citation network.

This analysis is similar to the previous one but with journal citations.



Main statistics about the network

 Size                                  2876 
 Density                               0.034 
 Transitivity                          0.364 
 Diameter                              4 
 Degree Centralization                 0.481 
 Average path length                   2.284 
 

4: Conceptual structure - Co-word analysis

Co-Word Networks show the conceptual structure, which uncovers the links between concepts through co-occurrences of terms.

The conceptual structure can be used to understand the topics being discussed (research front) and to identify which are the most important and most recent topics.

The tool divides the whole time span into different periods and compares the conceptual structures which is useful to analyze the evolution of the topics over time.

The package is able to analyze keywords, but also terms in article titles and abstracts. It does this by means of network analysis or correspondence analysis (CA) or multiple correspondence analysis (MCA). CA and MCA visualize the conceptual structure in a two-dimensional graph, which I also show below.

Joint word analysis using keyword co-occurrences

As with the previous graph, the options are to make the graph more readable:

  • normalize = "association" (the vertex similarities are normalized using association strength)

  • n = 50 (the function traces the top 50 cited references)

  • type = "fruchterman" (The network trace is generated using the Fruchterman-Reingold algorithm)

  • size.cex = TRUE (the size of the vertices is proportional to their degree)

  • size = 20 (maximum size of the vertices)

  • remove.multiple = FALSE (multiple edges are not removed)

  • labelsize = 3 (defines the maximum size of the vertices labels)

  • label.cex = TRUE (size of vertices labels is proportional to their degree)

  • edgesize = 10 (thickness of the edges is proportional to their degree. Edgesize defines the maximum value of the thickness)

  • label.n = 30 (labels are drawn only for the top 30 vertices)

  • edges.min = 25 (Only plot edges with a strength greater than or equal to 25)

  • all other arguments assume default values.

NetMatrix <- biblioNetwork(wos, analysis = "co-occurrences", network = "keywords", sep = ";")
net=networkPlot(NetMatrix, normalize="association", n = 50, Title = "Keyword Co-occurrences", type = "fruchterman", size.cex=TRUE, size=20, remove.multiple=F, edgesize = 10, labelsize=5,label.cex=TRUE,label.n=30,edges.min=2)

Joint word analysis by means of correspondence analysis

suppressWarnings(
CS <- conceptualStructure(wos, method="MCA", field="ID", minDegree=15, clust=5, stemming=FALSE, labelsize=15,documents=20)
)

5: Thematic maps

The co-word analysis draws clusters of the keywords. Themes are considered, whose density and centrality can be used to classify the themes and draw a two-dimensional diagram.

The thematic map is a very intuitive diagram and we can analyze the topics according to the quadrant in which they are placed: (1) upper right quadrant: main topics (as qualified by the authors of the tool); (2) lower right quadrant: basic topics; (3) lower left quadrant: emerging or disappearing topics; (4) upper left quadrant: very specialized/niche topics.

El análisis de co-palabras dibuja clusters de las palabras clave. Se consideran temas, cuya densidad y centralidad pueden utilizarse para clasificar los temas y trazar un diagrama bidimensional.

Map=thematicMap(wos, field = "ID", n = 250, minfreq = 4,
  stemming = FALSE, size = 0.7, n.labels=5, repel = TRUE)
plot(Map$map)

The description of the cluster can then be requested from the program:

Clusters=Map$words[order(Map$words$Cluster,-Map$words$Occurrences),]
library(dplyr)

Attaching package: 'dplyr'
The following objects are masked from 'package:stats':

    filter, lag
The following objects are masked from 'package:base':

    intersect, setdiff, setequal, union
CL <- Clusters %>% group_by(.data$Cluster_Label) %>% top_n(5, .data$Occurrences)
CL
# A tibble: 24 × 9
# Groups:   Cluster_Label [4]
   Occurrences Words           Cluster Color    Cluster_Label Cluster_Frequency btw_centrality clos_centrality pagerank_centrality
         <dbl> <chr>             <dbl> <chr>    <chr>                     <dbl>          <dbl>           <dbl>               <dbl>
 1           9 willingness           1 #E41A1C… willingness                 106           579.         0.00193             0.0112 
 2           7 acceptance            1 #E41A1C… willingness                 106           565.         0.00189             0.00872
 3           6 familiar              1 #E41A1C… willingness                 106           159.         0.00176             0.00660
 4           5 eating behavior       1 #E41A1C… willingness                 106           334.         0.00180             0.00572
 5           5 food neophobia        1 #E41A1C… willingness                 106           226.         0.00181             0.00480
 6          48 neophobia             2 #377EB8… neophobia                   416          4804.         0.00242             0.0438 
 7          29 behavior              2 #377EB8… neophobia                   416          4508.         0.00239             0.0271 
 8          21 neophilia             2 #377EB8… neophobia                   416          1163.         0.00205             0.0219 
 9          13 evolution             2 #377EB8… neophobia                   416          1001.         0.00204             0.0130 
10          12 exploration           2 #377EB8… neophobia                   416           936.         0.00204             0.0136 
# ℹ 14 more rows

6: Social structure - Collaboration analysis

This last section is also interesting. Collaborative networks show how authors, institutions (e.g., universities or departments) and countries relate to others in the field we are analyzing, in this case neophilia/neophobia in tourism. For example, the first figure below is a “Co-author network”. It uncovers regular study groups, hidden groups of scholars and key authors. The second figure is called an “Educational Collaboration Network” and uncovers relevant institutions in a specific research field and their relationships.

Author collaboration network

NetMatrix <- biblioNetwork(wos, analysis = "collaboration",  network = "authors", sep = ";")
net=networkPlot(NetMatrix,  n = 50, Title = "Author collaboration",type = "auto", size=10,size.cex=T,edgesize = 3,labelsize=1)

Educational collaboration network

NetMatrix <- biblioNetwork(wos, analysis = "collaboration",  network = "universities", sep = ";")
net=networkPlot(NetMatrix,  n = 50, Title = "Edu collaboration",type = "auto", size=4,size.cex=F,edgesize = 3,labelsize=1)

Collaboration network between countries

Finally, we can also obtain a graph of the collaboration network between countries, which can be useful to see between which countries there is more collaboration on the topic in question. For example, it can be seen that the country that has the most relationships is the United Kingdom. Spain, on the other hand, establishes collaborations with Japan, South Africa, Indonesia, Germany and China.

wos <- metaTagExtraction(wos, Field = "AU_CO", sep = ";")
NetMatrix <- biblioNetwork(wos, analysis = "collaboration",  network = "countries", sep = ";")
net=networkPlot(NetMatrix,  n = dim(NetMatrix)[1], Title = "Country collaboration",type = "circle", size=10,size.cex=T,edgesize = 1,labelsize=0.6, cluster="none")

Final comments

Apart from this library that works in R, the authors have created an application that does the same without programming, and there is even an additional tool, also useful. The only thing you need is to have the DB in the correct format (merged WoS and Scopus, if you want, etc…) because otherwise, it does not read it but if you do it right, the result is the same as we have shown here but without programming R practically (only to transform the files).

References

This work is done with the R package bibliometrix and on the basis proposed by Aria and Cuccurullo, adapted and modified for our research line.

Aria, M., & Cuccurullo, C. (2017). bibliometrix: An R-tool for comprehensive science mapping analysis, Journal of Informetrics, 11(4), pp 959-9753643, (https://www.bibliometrix.org).

Citation

BibTeX citation:
@online{caro2022,
  author = {Caro, José},
  title = {Scientific {Maps} {Analysis} with the {`R`} Package
    `Bibliometrix`},
  date = {2022-09-09},
  url = {https://carobarrera.com/blog/2022/09/bibliometric/},
  doi = {10.5281/zenodo.10079724},
  langid = {en}
}
For attribution, please cite this work as:
Caro, José. 2022. “Scientific Maps Analysis with the `R` Package `Bibliometrix`.” September 9, 2022. https://doi.org/10.5281/zenodo.10079724.