Visualization of bibliometrics showing citation networks, impact factors, publication trends, and discipline outputs with researchers interacting.

Problems of Bibliometric Analysis: Critical Evaluation and an Annotated List of Problems

Introduction: What Does This Document Cover?

Bibliometric analysis (BA for short) is a research method that examines published scientific literature using quantitative techniques in order to produce an overall picture of a field. It generates tables and maps showing how many articles have been published on a topic, which authors, institutions, or countries stand out, which concepts are mentioned together, and how the field has evolved over time. In recent years, this method has spread rapidly across nearly every scientific discipline.

However, this rapid spread has brought serious questions with it. Ole Ellegaard’s article in the journal Scientometrics focuses precisely on this point: do bibliometric analyses genuinely contribute to science, or are easily produced but shallow studies merely proliferating? Without denying the method’s strengths, the article systematically addresses the weaknesses and pitfalls that arise in practice.

This document was prepared to gather the problems discussed in scattered form throughout the article and turn them into a single, classified, and annotated list, one by one. Each problem is presented with (1) a brief definition, (2) a detailed explanation that an average reader can understand, and (3) a note indicating the sources on which the article relies. For the technical terms used in the text, a separate Glossary of Technical Terms is provided at the end of the document.

How should this document be read?

  • The problems are grouped under nine thematic chapters. The document can be read from start to finish, or you may skip directly to the chapter of interest.
  • Each problem is written as a standalone block; it is understandable even when read on its own.
  • The pale-blue framed “Basis in the article” notes indicate which researchers the relevant view is attributed to in the article.
  • Concepts marked with a star (★) are defined in the Glossary of Technical Terms at the end of the document.

Note: The explanations below are a restated (summarized and interpreted) version of the discussions in the source article; they are not verbatim quotations from it. The views belong to the article itself and to the works it cites.

Chapter 1 — Problems Relating to Research Design and Scientific Foundations

The problems in this chapter concern the most basic legitimacy of a bibliometric study: is there a genuine research question, or has the method become an end in itself? The article’s central concern arises from here.

Problem 1.1  ·  Absence of a research question: the method itself becoming the goal

A sound scientific study begins with a real question that needs answering. In many bibliometric articles, however, the starting point is not a question but the method itself: someone says, “let us do a bibliometric analysis on this topic,” runs the statistical tools, and turns the resulting tables into an article. In this case the analysis ceases to be a means of answering a question and becomes a goal in its own right. The result is often a study that offers a superficial snapshot of the field without providing any new insight.

The article stresses that bibliometric analysis is genuinely valuable in only two cases: either when an original and clear research question guides the study, or when the analysis is combined with other review methods such as a systematic review★. Otherwise, what emerges is not research worthy of publication but a simple statistical description.

Basis in the article: the core criticism in the editorial by Watson and Hayter (2024); also Haley (2023) and Lazarides et al. (2025).

Problem 1.2  ·  Lack of scientific maturity and a “research-level” depth

The number of bibliometric articles increases exponentially each year. This growth raises the question: “have a significant portion of these studies reached the maturity required to be published in a reputable scientific journal?” The article argues that many studies fail to clear this threshold and that they may need to be classified as statistical reports presenting an overview of a field rather than as research.

This does not mean the method is worthless. The problem is that, to qualify as a “scientific article,” a study must meet certain criteria: a clear research question, a systematic method, a theoretical foundation, and a rigorous assessment of the validity of the results. For analyses that do not meet these criteria, traditional peer-reviewed journals may not always be the most appropriate venue.

Basis in the article: Watson and Hayter (2024); Li et al. (2024, 2025); the discussion in Wang (2025).

Problem 1.3  ·  Superficial descriptive analyses and the “colorful diagram” illusion

Descriptive analyses produced with free software and eye-catching colorful charts can look impressive at first glance. Yet this visual richness is often not proportional to the depth of the insight offered. According to the article, such studies provide only “modest” insights; that is, a large number of charts and colors combine with a small amount of genuine knowledge.

The danger is that the persuasiveness of the visual conceals the shallowness of the content. A reader may look at a complex and aesthetically pleasing network map and assume that a deep analysis lies behind it, when what underlies it may be nothing more than automatically generated basic counts.

Basis in the article: Watson and Hayter (2024).

Problem 1.4  ·  Automated production does not confer scientific adequacy

Data-collection and counting software automatically produces the most basic part of the analysis (number of publications, most productive author, most-cited country, and so on). The point the article underscores is this: output obtained by pressing a button in a program does not, on its own, justify a scientific publication. For here the researcher’s original contribution is negligible; anyone could produce the same result from the same data.

For scientific value to emerge, more advanced methods, original interpretations, and new syntheses must be added on top of this automated foundation. A study should go beyond updating existing counts or presenting an output reproducible with automated tools; it should improve the methods applied or combine prior knowledge in a new way.

Basis in the article: assessment of tools such as InCites / SciVal; Hulland (2024).

Chapter 2 — Problems Relating to Author Behavior and Publication Motives

This chapter looks at the human and institutional dimension of the problem rather than the technical one: why do people produce so many bibliometric analyses? The article points out that some of these motives stem from practical advantages rather than scientific curiosity.

Problem 2.1  ·  The phenomenon of “transient authors”

There are many authors who enter the field of bibliometrics from outside, produce only a single bibliometric article, and then leave the field. These are called “transient authors”★. Such people typically analyze a topic from their own area of expertise once, using easily accessible software, and then move on.

This situation has two drawbacks. First, it reflects an opportunistic approach: a topic is chosen not because there is a genuine research need but because the software is readily available. Second, it weakens the cumulative progress of the field, because these authors do not form a deep connection with the bibliometric literature and its methodological debates. The article recommends that newcomers to the field pay greater attention to the older and more comprehensive literature that discusses the method.

Basis in the article: González-Alcaide (2021); Watson and Hayter (2024).

Problem 2.2  ·  Strategic and career-oriented motives

For some researchers, the production of bibliometric analyses is driven by strategic and career-related calculations rather than a scientific aim. Because bibliometric articles can be produced relatively quickly and often attract high citation counts, they can become a practical way of expanding one’s publication list. This is an especially attractive option in academic environments where the “publish or perish” pressure is intense.

The problem lies less in the motive itself than in the way this motive pushes the quality of the study into the background. When the goal is to increase the number of publications rather than to answer a question, a proliferation of superficial analyses becomes inevitable.

Basis in the article: Ansorge (2024); Watson and Hayter (2024).

Problem 2.3  ·  The appeal of high citations: the review-article effect

Bibliometric analyses can be regarded as a kind of review article, and it is well known that review articles attract more citations than original research articles. This is an additional factor that makes them attractive to produce: with relatively little effort, one can obtain an output with a high citation return.

This appeal creates an incentive structure that puts quantity ahead of quality. Because its potential to attract citations is high, a study may be deemed valuable regardless of the depth of its content; yet high citations alone do not guarantee scientific quality.

Basis in the article: Cheng et al. (2024); Montazerian and Dorch (2022).

Problem 2.4  ·  AI tools lowering the methodological threshold

New AI-based tools such as SciSpace, Scite AI, or Perplexity have considerably eased data-organization and writing processes. Guides offering a “recipe” for bibliometric analysis, along with ready-made templates, have also markedly reduced the knowledge barrier to entry. While this is a positive development in itself, it carries a side effect.

As the method becomes easier to apply, it also becomes easier to produce analyses without genuinely understanding the underlying criteria. The article warns that this may create a new challenge for publishers: AI tools may cause a growing number of analyses with weak methodological awareness to flow into journals.

Basis in the article: Ahmi (2022); Donthu et al. (2021); Ashley (2025); Watson and Hayter (2024).

Problem 2.5  ·  Low-budget mass production and loss of reliability

Some of the criticisms point to the fact that bibliometric studies are often produced cheaply and in bulk, as if coming off an assembly line. The easy availability of cheap (or even free) tools sets the stage for possible overproduction.

This mass production can undermine the reliability and impact of the studies, for speed and low cost often come at the expense of methodological rigor and in-depth interpretation. As a result, the literature fills up with a large number of studies that are weak in content.

Basis in the article: Li et al. (2025); Repiso and Cabezas-Clavijo (2025).

Chapter 3 — Problems of Methodological Standardization and Transparency

Systematic reviews have established reporting guidelines such as PRISMA★; everyone writes and reads according to the same rules. In bibliometric analyses, however, such a shared standard was long lacking. This chapter addresses the problems created by that gap.

Problem 3.1  ·  Lack of a common guideline and standard

For a long time there was no agreed-upon common set of rules for how bibliometric analyses should be conducted and reported. While guides such as PRISMA★ have been used for systematic reviews for years, comparable guidance for bibliometrics has been scattered across the literature and the number of evidence-based studies has remained limited. This makes it difficult to compare studies with one another and to reproduce them.

The article introduces initiatives that try to fill this gap: BIBLIO★ for the biomedical field, the general-purpose 32-item GLOBAL★, PRIBA★ for the health-medicine field, and the reader/reviewer-oriented VALOR★ framework. The common aim of these guidelines is to provide more reliable, standardized, and reproducible analyses; however, a single universal standard has not yet taken hold.

Basis in the article: Cheng et al. (2024); Fei and Hu (2024); Montazeri et al. (2023, BIBLIO); Ng et al. (2025, GLOBAL); Koo and Lin (2023, PRIBA); Hoang (2025, VALOR).

Problem 3.2  ·  Inadequate explanation of method selection

Researchers often fail to adequately explain which bibliometric method they chose and why, and what new insight that choice provides. Yet bibliometrics rests on statistical methods, and each method has its own assumptions, limits, and sources of error. When these are not stated, the reader cannot assess how reliable the results are.

Transparency is the key concept here. A study should set out, as clearly as possible, the methods used, the limitations, and the possible sources of error. Otherwise the results are left hanging, because the foundation on which they rest is not visible.

Basis in the article: Haley (2023); Lazarides et al. (2025).

Problem 3.3  ·  Unawareness of methodological limitations

A significant proportion of authors are insufficiently aware of the limits of the methods they use, of the potential lack of depth, and of the sources of error. Bibliometrics is statistically grounded, but the rules of statistics are not always applied with awareness. This lack of awareness leads to misinterpretations and overgeneralizations.

The article stresses that the most important solution to this is training and collaboration: newcomers to the field, when using advanced measurement and mapping techniques that often require knowledge beyond self-taught skills, should work more systematically with trained bibliometricians. Courses, webinars, and workshops offered by centers such as CWTS (Leiden) can help close this gap.

Basis in the article: Glänzel (1996); Wallin (2005); Thompson and Walker (2015); Torres-Salinas et al. (2025).

Problem 3.4  ·  The reproducibility problem

For a scientific result to be valuable, others must be able to follow the same steps and arrive at the same result. In bibliometrics this is often neglected. Because databases are continually updated, the date on which the data were downloaded must always be stated; otherwise the same search may return a different result a month later and the study becomes impossible to repeat.

Likewise, all parameter settings, clustering options, and link-pruning★ decisions used must be clearly documented. If the study is later to be updated, expanded, or carried over to other fields, it is impossible to proceed without these records. In short: reproducibility is not a luxury added afterward but a necessity that must be designed in from the very start.

Basis in the article: Fei and Hu (2024).

Problem 3.5  ·  Methodological diversity hindering comparison and collaboration

Bibliometricians use different approaches, different software, and different criteria. Although this diversity may look like richness at first glance, in practice it creates a problem: the results of analyses performed with different methods become incomparable, and collaboration among researchers becomes harder.

The article also notes that there is weak collaboration and missing citation links between the bibliometrics community and the experts of the outside fields (for example, medicine, engineering) that use the method. When those who know the method and those who know the subject do not come together sufficiently, analyses remain deficient both methodologically and substantively.

Basis in the article: Fei and Hu (2024); González-Alcaide (2021).

Chapter 4 — Problems Relating to Data Sources and Databases

A bibliometric analysis is only as good as the data that feed it. This chapter addresses the problems arising from the choice and limitations of databases such as WoS★, Scopus★, and Google Scholar★.

Problem 4.1  ·  Database limitations and bias seeping into the results

Which database is chosen directly affects the outcome of the analysis. If a database under-covers certain journals, languages, or regions, that gap is reflected in the conclusions as a bias. Similarly, choosing too narrow a time window also distorts the picture.

The article recommends giving priority to indexing quality and data compatibility when selecting a database. The question is: do the databases used contain enough sources, or would broadening the range of data strengthen the analysis? Relying on a single database is practical, but it carries the risk of presenting an incomplete picture.

Basis in the article: Harzing and Alakangas (2016).

Problem 4.2  ·  Multiple indexing: the same publication being counted more than once

Databases such as WoS★ and Scopus★ may index a publication under more than one subject category; likewise an article may be linked to several institutions or countries. This “multiple indexing” distorts statistical counts: when a publication is counted again and again across different categories, the true weight of fields or countries may appear different from what it actually is.

This is, on the surface, a technical detail; but it seriously affects the interpretation of results. When a researcher, unaware of this, says “this field has produced this many publications,” they may in fact be reporting an inflated figure.

Basis in the article: assessment of the WoS / Scopus indexing structure (the article’s own observation).

Problem 4.3  ·  Neglect of specialized databases (the data-integration problem)

WoS★ and Scopus★ dominate standard analyses; however, field-specific databases such as Medline (medicine), Inspec (physics), or PsycInfo (psychology) can cover a topic in much greater depth. These specialized sources can enrich the collected data with field-specific references.

The problem is that researchers often ignore these databases. The reason is that different databases use different metadata standards, which makes bringing the data together (integration) difficult. For the sake of convenience, field-specific depth is given up.

Basis in the article: assessment of data integration and metadata differences (the article’s own observation).

Problem 4.4  ·  Google Scholar’s superior coverage but technical impracticality

Google Scholar★ (GS) generally indexes more than twice as many references in a field as WoS★ or Scopus★; this includes gray literature★, reports, articles in small journals, and non-English papers. This is valuable coverage for a complete picture, and in most cases GS also records far more citations than the other databases.

Yet this richness cannot be exploited easily. Google Scholar has no API★ (an interface for automated data retrieval), and the bulk downloading of references is extremely limited. For this reason GS data cannot readily be incorporated into a bibliometric workflow. Analyses based solely on WoS or Scopus, meanwhile, miss many of the citations that GS captures.

Basis in the article: Moral-Muñoz et al. (2020); Harzing and Alakangas (2016).

Problem 4.5  ·  “Dark citations”

Citations that do not enter traditional databases and appear in sources outside regular peer-reviewed journals are called “dark citations”★. For example, in the public-health literature, citations in federal sources produced by government agencies fall into this category and are not seen by standard databases.

This is an important blind spot of citation analysis. The true impact of a study may appear smaller than it is when measured only by citations in peer-reviewed journals, because its impact in policy documents, guidelines, or institutional reports is not taken into account.

Basis in the article: Keralis et al. (2023).

Problem 4.6  ·  Author-identity ambiguity

Different authors with the same name, inconsistencies in name spelling, changing addresses, and institutional affiliations threaten the accuracy of bibliometric data. If there are ten different researchers named “M. Yılmaz,” their contributions may be mistakenly attributed to a single person, or vice versa.

To reduce this problem, the article recommends author-identification tools: ORCID★, ResearcherID (WoS), and ScopusID. These identifiers should be part of the effort to clean up duplicate records and “clean” the data. Otherwise, even the most basic counts become unreliable.

Basis in the article: Jonkers and Derrick (2012).

Chapter 5 — Problems of Search Strategy and Data Cleaning

Every bibliometric analysis begins with a search query. This chapter addresses the problems that show how much more critical the question “which articles should we include?” is than it appears. A poorly designed search invalidates every subsequent step.

Problem 5.1  ·  The precision–recall balance

A search query has two basic measures. Precision★ indicates how much of what is found is genuinely relevant to the topic, while recall★ indicates how much of all the relevant articles are captured. The ideal is for both to be high; but in practice, as one rises the other tends to fall.

A very broad query captures everything relevant (high recall) but also brings in plenty of irrelevant results (low precision). A very narrow query does the opposite. The article recommends that search profiles and search fields be evaluated and documented with this balance in mind. An iterative (trial-and-improvement) search strategy ensures that fewer relevant articles are missed.

Basis in the article: Jonkers and Derrick (2012); Lim et al. (2024).

Problem 5.2  ·  Overly broad keywords and “false” results

Keywords formulated too broadly cause irrelevant “false” results to enter the data set. These false results must later be weeded out; otherwise they distort the results of the analysis. Including both too many and too few articles has the potential to distort the inferences.

According to the article, the only way to fully control this problem is the manual review and selection of references by hand. The ideal method is for two researchers to evaluate the articles independently and include only those they agree are relevant. This is time-consuming with large data sets; however, tools such as Covidence and AI techniques can ease the process.

Basis in the article: Lim et al. (2024); Covidence and AI-assisted screening approaches.

Problem 5.3  ·  The choice of search fields (ti, kw, ab, ts)

In databases, a search can be carried out in different fields: title (ti)★, keyword (kw)★, abstract (ab)★, or a broad topic (ts)★ field covering all of these. Which of these fields is used, and whether they are used individually or together, determines which articles are captured and significantly changes the result.

The problem is that this choice is often made without thought and is not reported. Searching only in the title gives a more precise but more incomplete result; searching in the broad topic field is more inclusive but noisier. The researcher should make this choice deliberately and explain the rationale for it.

Basis in the article: assessment of the choice of search fields (the article’s own observation).

Chapter 6 — Problems Relating to Software Tools and Visualization

Bibliometric analysis depends heavily on software such as VOSviewer★ and CiteSpace★. This chapter addresses both the strengths and the misleading aspects of these tools. One of the article’s central concerns lies here: using tools is not the same thing as understanding them.

Problem 6.1  ·  Over-reliance on software and using metrics without understanding them

The most fundamental problem, highlighted in the article’s abstract as well, is that researchers become over-reliant on software without genuinely understanding the underlying metrics. A program produces a number or a map; but if it is not known how that number is calculated, on what assumptions it rests, and what is overlooked, the output is misinterpreted.

Bibliometrics uses advanced measurement and mapping techniques that go beyond simple counts; these often require knowledge beyond self-taught skills. Running the software without knowing “what it does” makes the result unreliable.

Basis in the article: Ellegaard (2026) abstract and introduction; Glänzel (1996); Thompson and Walker (2015).

Problem 6.2  ·  Different tools producing different results from the same data

Different software uses different techniques and different visualizations. The striking consequence is this: the same data set, when given to different tools, can lead to different interpretations. That is, the result depends partly not on the data but on the tool chosen.

This shows that bibliometric results are not as “objective” as they appear. To ensure reproducibility, the article recommends that the tool used and all parameter settings always be disclosed. Where possible, the results should be compared using different tools or simpler approaches.

Basis in the article: Hoang (2025); Chen et al. (2022); Fei and Hu (2024).

Problem 6.3  ·  Confusion between full counting and fractional counting

When counting contributions, software uses one of two methods. Full counting★ counts an article in full for each of its authors/institutions; fractional counting★ divides the contribution among the authors. Which one is chosen determines how much weight multi-authored articles or review articles carry within the network.

This counting process often creates confusion. For example, the Bibliometrix★ tool counts a document only once when examining author institutions but counts each department separately. If the researcher is not aware of this difference, they will misinterpret the results. The article stresses that users must clearly state which counting method they use.

Basis in the article: van Eck and Waltman (2014); Aria and Cuccurullo (2017).

Problem 6.4  ·  CiteSpace’s steep learning curve and complexity

CiteSpace★ produces powerful, dynamic, graph-based visualizations showing the development of a field of knowledge over time; it effectively takes an “X-ray” of a field. But the price of this power is high: according to the article, its most important disadvantage is the steep learning curve required to set the sensitive visualization parameters.

The maps and clusters the program produces can be complex, and interpreting them requires specialized domain knowledge. The presence of so many options makes the interface overwhelming for inexperienced users. This is a serious problem especially for the “transient authors”★ who are new to the field and perform only a single analysis, because they use the tool without fully learning it.

Basis in the article: Chen (2006, 2016); González-Alcaide (2021).

Problem 6.5  ·  “Mysterious-looking” visualizations

The article relays a striking observation by Chen (2016), the developer of CiteSpace★: producing mysterious-looking visualizations with such tools is often easier than fully understanding what these visuals show and who would find them useful. That is, producing an aesthetic complexity is simpler than making sense of it.

Combined with the “colorful diagram illusion” in Section 1.3, this creates a serious risk: the more impressive the visual, the more the viewer assumes that some deep meaning lies there. Yet complexity is often a sign not of depth but of a lack of interpretation. Animated models can show the features of a field; but they are cognitively extremely tiring.

Basis in the article: Chen (2016); Synnestvedt et al. (2005).

Problem 6.6  ·  VOSviewer’s limited analytical capacity

Thanks to its intuitive interface, VOSviewer★ is the most widely used tool in standard analyses and is strong at clustering. However, as its developers also note, it does not include some of the metrics needed for more advanced analysis: metrics such as betweenness centrality★, silhouette value★, modularity★, and burst detection★ are absent in VOSviewer; these are found in CiteSpace★. Moreover, in VOSviewer it is not possible to split the results into time series.

Because of its strong focus on visualization, the article relays that VOSviewer offers less functionality than other tools for analyzing bibliometric networks. That is, the tool is successful at displaying networks but limited at in-depth analysis. Without knowing this limit, a user may suppose the tool to be suitable for every purpose.

Basis in the article: van Eck and Waltman (2014); Eck and Waltman (2014).

Problem 6.7  ·  Insufficient understanding of program capabilities

Another fundamental problem is uncertainty about what analysis programs are capable of. Users often do not know the full capacity of the tool. This has two consequences: the data are not used adequately (the advanced options the program offers are skipped) and interpretations become biased.

The article characterizes this as a clear analytical challenge. Mastering a tool’s capabilities is a precondition both for extracting the most information from the data and for making a balanced interpretation. An analysis performed with a half-learned tool produces a picture that is both incomplete and skewed.

Basis in the article: Hulland (2024).

Chapter 7 — Problems Relating to Clustering and Parameter Selection

This chapter is the technical problem area the article dwells on most. Clustering★ is the process of dividing articles into groups according to their shared features, and it lies at the heart of bibliometric maps. Yet these groups are extremely sensitive to the parameters the researcher selects; a small change in a setting can produce an entirely different “map of the field.”

Problem 7.1  ·  The vagueness of clustering criteria

In the bibliometric literature, very few studies clearly state the criteria by which clustering★ was performed. The clusters produced by algorithms necessarily require manual (human) interpretation; but the basis on which this interpretation rests often remains in the dark. The reader cannot understand why the clusters formed as they did.

This vagueness directly undermines the reliability of the results. When a “map of the field” is presented without showing which decisions produced that map, the map turns from scientific evidence into a design choice.

Basis in the article: assessment regarding the non-reporting of clustering criteria; van Eck and Waltman (2023).

Problem 7.2  ·  Synonym management

How do programs handle different terms that mean the same thing (for example, “artificial intelligence” and “AI”)? This “synonym management” seriously affects the structure of the clusters. The standard method is to use a thesaurus★ (a list of concept equivalences) that matches terms; but how this list is built and which terms are merged is often not explained.

The way synonyms are merged can change the number, size, and content of the clusters. If “AI” and “artificial intelligence” are counted separately, this concept is artificially split into two small clusters; if merged, a single, strong cluster forms. When this decision is not documented, the map takes on an arbitrary appearance.

Basis in the article: Tan et al. (2025); van Eck and Waltman (2023).

Problem 7.3  ·  The effect of cluster number and size on the results

How many clusters form in an analysis and how large these clusters are depend directly on the selected parameters. The article notes that studies rarely examine the extent to which the number and size of clusters affect bibliometric results. Yet this can fundamentally change the outcome.

The article gives a striking hypothetical example: when a specialty area is divided into a small cluster, a temporal analysis may show marked growth in that field. When larger clusters are used on the same data, the same growth appears insignificant. That is, even the answer to the question “is the field growing or not?” can change depending on the chosen cluster size.

Basis in the article: van Eck and Waltman (2023).

Problem 7.4  ·  Wrong parameter choice: fragmentation and missed trends

Choosing the parameters in the programs incorrectly leads to errors in two directions. On one hand, the network can become excessively fragmented (topics that are in fact related fall into artificially separate clusters); on the other hand, real trends can be missed (an important development becomes invisible because of a wrong setting).

The user also makes decisions such as links pruning★ and cluster merging on their own. Using these automatic algorithms requires establishing some fixed reference points to be sure one is on the right track. Without these points, the researcher may mistake a wrong map for a correct one.

Basis in the article: assessment regarding parameter selection; van Eck and Waltman (2023).

Problem 7.5  ·  The resolution parameter and the vagueness of “quality”

In tools such as VOSviewer★ there is a resolution parameter★ that determines how fine or coarse the clusters will be. The developers’ advice is surprisingly open-ended: the user is advised to try different resolution values and choose the one that gives the level of detail best suited to their purpose. But here what “suitable” or “of quality” means is not clearly defined.

The deep conclusion the article underscores is this: in this case clustering ceases to be an operation leading to a single, definite result; it turns into a parallel interpretation process producing solutions at different levels of detail. This means that one moves partly away from the view that bibliometric mapping is a purely quantitative method, and that the analysis in fact also rests on qualitative judgments.

Basis in the article: van Eck and Waltman (2023).

Problem 7.6  ·  Threshold sensitivity: noise or blindness?

In an analysis there are threshold values★ that determine how many co-citations★ or keywords will be taken into account. This threshold determines the overall size of the network and how the clusters will form. A trade-off (balance) is at work here.

Low threshold values produce dense but “noisier” networks (irrelevant links fill the picture). High threshold values create the risk that genuine links are missed (the network is simplified but important relationships are lost). The right threshold must be chosen carefully according to the field and purpose and must be reported; otherwise the result is either drowned in noise or rendered blind.

Basis in the article: assessment regarding the choice of threshold values (the article’s own observation).

Problem 7.7  ·  The effect of time-window selection

As the literature is divided into time windows, the analysis becomes harder, and the size of the selected time windows changes the results. Especially when identifying “research fronts” (a field’s most current, most active topics), which time intervals are chosen is decisive. Different time windows can show different fronts.

For this reason the article recommends interpreting time-based results cautiously: these results should be evaluated so that they coincide with the results of simpler approaches or evidence from other sources. That is, the output of a temporal analysis should be considered not on its own but together with other evidence.

Basis in the article: Chen et al. (2022).

Problem 7.8  ·  Inevitable information loss in visualization

Bibliometric networks make complex information understandable by simplifying it; this is their strength. But as the article stresses, this simplification has a cost: information loss. When bibliographic data are reduced to nodes (points), some information is always lost.

For example, when text data are converted into a term co-occurrence network★, the information about the context in which the terms appear together is lost. Likewise, in a citation network one can see who cites whom, but one cannot see why a citation was made. Many authors are unaware of the information loss caused by this reduction, which is performed to make the diagrams more interpretable.

Basis in the article: Van Eck and Waltman (2014); Eck and Waltman (2014).

Problem 7.9  ·  Quantitative or qualitative? The method’s identity dilemma

Bibliometric mapping is usually presented as a quantitative (numerical) method. However, as seen in Section 7.5, the fact that parameter choices are subjective and that the “most suitable” solution is determined according to the researcher’s purpose shows that the method in fact contains a significant amount of qualitative (interpretive) judgment.

This dilemma is important for the method’s legitimacy. If the results depend partly on the researcher’s choices, bibliometric analysis cannot fully sustain its claim to “objective counting.” For this reason, the article recommends cross-validating the results (triangulation★) with methods such as systematic review★ or qualitative assessment.

Basis in the article: van Eck and Waltman (2023); the article’s synthesis.

Problem 7.10  ·  Method–result inconsistency: same data, different themes

The article illustrates a striking point with a concrete example. When the same 2023 data are examined by two different methods (one based on WoS★ journal categories, the other on the co-occurrence of keywords★), no clear agreement appears in the main themes that emerge. The general patterns are recognizable; but the division of the themes changes according to the method.

Likewise, in a co-citation★ analysis performed with CiteSpace★, the cluster themes only partly overlap with the themes of the other methods. The meaning is clear: the division into clusters and subject areas depends largely on the chosen analysis method. This shows why the assumptions and method choices need to be clearly explained.

Basis in the article: comparison of Figure 2, Table 1, Figure 3, and Figure 4 in the article.

Chapter 8 — Problems Relating to Citation Analysis and the Use of Metrics

Citation count is bibliometrics’ most common measure of “impact”; but it is also the most contested. This chapter addresses the conceptual problems that arise from using citations as a measure.

Problem 8.1  ·  Over-reliance on citation-based metrics

Relying solely on citation analysis and citation-based metrics to evaluate researchers or institutions is a practice that has long been criticized in the literature. This concern has become so widespread that a group of editors and publishers launched an initiative such as DORA★ (the Declaration on Research Assessment), which advocates halting the use of citation-based metrics in evaluating researchers.

The article stresses that citation analysis is not sufficiently multidimensional and that, on its own, it remains inadequate for reflecting complex contexts. For this reason it recommends that bibliometric evaluations be based on different metrics. Initiatives such as the Leiden Manifesto★ and CoARA★ likewise support a more diversified and responsible use of metrics.

Basis in the article: American Society for Cell Biology (2012, DORA); Hicks et al. (2015, Leiden Manifesto); CoARA (2023); Moed (2020); Haley (2023).

Problem 8.2  ·  The “citation = quality” fallacy

A common but contested assumption is to equate quality with citation count. According to the article, this view is frequently debated, because citations do not have the same “value” and cannot be counted with equal weight. A paper receiving many citations does not necessarily make it good.

The value of citations varies greatly according to the context in which they arise and the purpose for which they are given. A citation may have been given to confirm a claim, or merely to list the existing literature. These two are counted as the same quantitatively; but their scientific meanings are diametrically opposed.

Basis in the article: Haley (2023); Moed (2020).

Problem 8.3  ·  Disregard for citation context and motive

Traditional citation counting cannot see why a citation was made; it only counts that it was made. Yet a study may have been cited as a positive reference, in order to be criticized, or merely out of courtesy. When context is ignored, citation count becomes a misleading measure of impact.

The article introduces new AI tools that reduce this problem: programs such as Scite AI analyze the context in which a paper is cited and weight the citations. Similarly, the “Automated Research Impact Assessment” developed by NIEHS draws on citations in “important” sources such as policy documents, regulations, and clinical guidelines. These approaches represent a shift from the “quantity” of citations to their “quality.”

Basis in the article: Ashley (2025); Drew et al. (2016).

Problem 8.4  ·  Cross-field citation differences and old-literature bias

Citation practices vary greatly from field to field. While the natural sciences and medicine use citation heavily, the situation is the opposite in the humanities. For this reason it is misleading to compare citation counts from different fields directly; a citation count considered “low” in one field may be “high” in another.

In addition, citation count naturally favors older literature, because older articles have had more time to accumulate citations. Review articles and self-citation can also artificially inflate the counts. To reduce these problems, the article recommends applying normalization★ methods (making data of different origins comparable); this concerns not only basic citation analysis but also all derived indicators such as the h-index★, co-citation★, and bibliographic coupling★.

Basis in the article: Haustein and Larivière (2015).

Problem 8.5  ·  Neglect of content analysis (peer review)

In research evaluation, far less importance is given to content analysis (that is, expert peer review) than to quantitative indicators. Numbers stand out because they are easy and fast; yet understanding the true value of a study requires its content to be read by experts.

The article advocates highlighting qualitative approaches for a balanced evaluation. DORA★ likewise recommends shifting the focus to the content of documents and clearly stating the evaluation criteria. Numbers should at most be a complement to expert judgment, not a replacement for it.

Basis in the article: Haustein and Larivière (2015); American Society for Cell Biology (2012, DORA); Montazerian and Dorch (2022).

Chapter 9 — Problems Relating to Publication and Dissemination

The final chapter looks at the “where should it be published?” dimension of the problem. The article’s most provocative question is here: does every bibliometric analysis truly deserve to be published in a peer-reviewed journal?

Problem 9.1  ·  The “gray journals” problem

A large share of bibliometric analyses in the health sciences are published by publishers referred to as “gray journals”★ (for example, Frontiers or MDPI). This situation has triggered a debate about the quality and impact of these studies.

According to the criticism, these studies are often produced cheaply and en masse, which can undermine their reliability and impact. Moreover, although more bibliometric analyses than ever are being published in the health sciences (especially from China), it has been noted that their impact remains limited and that more international collaboration is needed in the field.

Basis in the article: Wang (2025); Li et al. (2024, 2025).

Problem 9.2  ·  Unsuitability of the publication channel

One of the article’s boldest suggestions is this: for a narrowly focused, descriptive bibliometric analysis, a traditional peer-reviewed journal may not always be the most suitable venue. In some cases, being a study report, a memo, or part of an institutional/annual report may be a more fitting option.

This suggestion does not mean the study is worthless; on the contrary, it says the study should reach the right audience through the right channel. Policy briefs, institutional evaluation reports, and grant applications already use bibliometric analysis increasingly, often in non-traditional forms. The article does not pass a definitive judgment on this; but it invites the bibliometrics community and editors to debate the issue.

Basis in the article: Watson and Hayter (2024); Kurutkan (2026).

Problem 9.3  ·  The short lifespan of descriptive analyses

Bibliometric analyses, especially those based on citation analysis and productivity data, risk being short-lived by their very nature. This is because collaboration patterns and research areas change over time; a map valid today may be outdated a few years later.

For this reason, simple, descriptive analyses often offer only a “snapshot” of a field and remain valid for only a short, limited period. To stay current as scientific documents, they need to be continually renewed; this makes them rather unsuitable for traditional journal publishing. This limitation makes it all the more important that studies clearly state their assumptions, limits, and statistical uncertainties.

Basis in the article: the article’s conclusion section; Wallin (2005).

Problem 9.4  ·  Lack of formal bibliometric training

Formal bibliometrics training at universities is almost nonexistent. This shortfall creates a body of users who learn the method on their own, often only partially. This is one of the root causes of the problems listed throughout many sections of this document (wrong parameter choice, using metrics without understanding them, being unaware of limitations).

To close this gap, the article suggests concrete avenues: online courses, webinars, and professional workshops from centers such as CWTS (Leiden); and participation in research-evaluation reform movements such as DORA★ and CoARA★. These initiatives both build skills and provide platforms for networking and standard-setting for the bibliometrics community. The ultimate aim is to increase transparency and reproducibility.

Basis in the article: González-Alcaide (2021); Torres-Salinas et al. (2025); Ng et al. (2025).

Collective View of the Problems (Synthesis Table)

The table below was prepared to see all the problems addressed in detail in the document at a single glance. Each row summarizes one problem, the chapter it belongs to, and its essence in a single sentence.

No.ThemeEssence of the problem
1.1DesignThe method itself becoming the goal instead of a genuine research question.
1.2DesignMost studies failing to reach the scientific maturity that would merit publication.
1.3DesignColorful charts creating an illusion of depth that masks shallow content.
1.4DesignAutomatically generated basic counts not justifying a publication on their own.
2.1MotivesThe proliferation of “transient authors” who enter and leave the field with a single article.
2.2MotivesMotives focused on career and publication count rather than a scientific aim.
2.3MotivesHigh citation potential putting quantity ahead of quality.
2.4MotivesAI tools lowering the methodological threshold and easing shallow production.
2.5MotivesLow-budget mass production undermining reliability.
3.1StandardizationThe long absence of an agreed common reporting guideline.
3.2StandardizationThe rationale for the method choice not being adequately explained.
3.3StandardizationLack of awareness of the method’s limitations and sources of error.
3.4StandardizationResults not being reproducible; the download date not being stated.
3.5StandardizationMethod diversity hindering comparison and collaboration.
4.1DataDatabase limitations seeping into the results as bias.
4.2DataThe same publication being counted multiple times due to multiple indexing.
4.3DataField-specific databases being overlooked because of integration difficulty.
4.4DataGoogle Scholar being unusable due to the lack of an API despite its broad coverage.
4.5Data“Dark citations” that do not enter traditional databases being missed.
4.6DataAuthor-name ambiguity distorting the counts.
5.1SearchThe balance between precision and recall not being observed.
5.2SearchOverly broad keywords contaminating the data with irrelevant results.
5.3SearchSearch fields (title/abstract/topic) being chosen without deliberation.
6.1SoftwareOver-reliance on software without understanding the metrics.
6.2SoftwareDifferent tools producing different results from the same data.
6.3SoftwareThe confusion created by the full/fractional counting difference.
6.4SoftwareCiteSpace’s steep learning curve and complexity.
6.5Software“Mysterious-looking” visuals being easier to produce than to interpret.
6.6SoftwareVOSviewer lacking advanced analysis metrics.
6.7SoftwareProgram capabilities not being adequately understood; incomplete use of the data.
7.1ClusteringClustering criteria being hardly ever explained.
7.2ClusteringThe handling of synonymous terms silently altering the clusters.
7.3ClusteringThe number and size of clusters being able to reverse the results (e.g., growth).
7.4ClusteringWrong parameters leading to fragmentation and missed trends.
7.5Clustering“Quality” in the resolution parameter remaining undefined and subjective.
7.6ClusteringThe threshold value causing either noise or blindness.
7.7ClusteringThe choice of time window changing the research fronts.
7.8ClusteringReduction to a network creating inevitable information loss.
7.9ClusteringThe method appearing quantitative while actually resting on qualitative choices.
7.10ClusteringThe same data producing different themes with different methods.
8.1CitationOver-reliance on citation metrics alone in evaluation.
8.2CitationThe fallacy of equating quality with citation count.
8.3CitationDisregarding why a citation was made (context and motive).
8.4CitationIgnoring cross-field citation differences and old-literature bias.
8.5CitationContent/peer review being neglected in favor of quantitative indicators.
9.1PublicationLow-impact mass production in “gray journals.”
9.2PublicationPeer-reviewed journals not being the suitable venue for every narrow analysis.
9.3PublicationThe short-lived, “snapshot” nature of descriptive analyses.
9.4PublicationFormal bibliometrics training being almost nonexistent at universities.

Solutions and Minimum Criteria Proposed by the Article

The article does not merely list the problems; it also offers concrete suggestions for overcoming them. The minimum criteria and general recommendations set out below by the article — as a complement to established guidelines such as BIBLIO★ and GLOBAL★ — are summarized here.

Core criteria for an analysis to be considered scientific

  1. The study must pose a clear research question or problem that needs to be clarified.
  2. It must rest on systematic methods, most often grounded in empirical data and a theoretical foundation.
  3. It must take a reflexive (self-critical) approach toward both theory and method; in particular, all assumptions and parameters in the clustering techniques used in mapping must be carefully explained.
  4. The reasoning and the validity behind the results must be carefully evaluated.

Minimum reporting criteria to be met before publication

  • Clear units of analysis must be defined; the study’s potential meaning for the target audience must be stated.
  • The precision★ and recall★ of the search set must be documented together with the chosen database and search profile.
  • Data cleaning and reduction operations must be clearly stated.
  • In citation analyses, the dependence of the results on the chosen time window and databases, as well as the limits of using citation count as a measure of impact, must be stated.
  • All software tools and parameter settings used must be explained; clustering procedures must be reported fully and transparently. Advanced metrics that not everyone may be familiar with must be defined.
  • The use of clustering and link pruning★ in the networks, and their results, must be stated; the results must be fully reproducible.
  • The possibility of using alternative metrics★ (altmetrics) for the present analysis must be discussed.
  • Especially in macro analyses, whether experts in the field support the inferences must be stated.

Additional recommendations for authors, editors, and reviewers

  • The study must offer a new contribution (improving the method or synthesizing prior knowledge in a new way); merely updating old counts or being reproducible with automated tools is not enough.
  • Authors must state the purpose and target audience of their work and focus on producing contextualized insight rather than numbers and charts.
  • As part of the process, experts in the field being analyzed must be consulted. The article goes further and recommends that at least one expert from the relevant field take part in the peer review of bibliometric studies, and even that editors encourage co-authorship by domain experts.
  • Triangulation★ is recommended: bibliometric findings based on large clusters must be corroborated with a systematic review★ or qualitative evaluation.
  • Where appropriate, alternative publication channels such as policy briefs, institutional reports, or working notes must be considered for narrowly focused analyses.

In short, the article’s core message is this: bibliometric analysis is a valuable tool; but it makes a genuine contribution to science only when it is guided by a clear question, when its method and limits are explained transparently, when it is combined with domain expertise, and when it is disseminated through the right channel. The aim is fewer but more valuable publications.

Glossary of Technical Terms

Below are explanations, aimed at the average reader, of the technical terms frequently used in bibliometrics that are marked with a star (★) in the document. The terms are grouped according to their related topics.

A. Basic Concepts and Methods

TermDefinition
Bibliometric analysis (BA)A research method that reveals the structure, trends, and actors of a field by examining published literature with numerical/statistical techniques.
Performance analysisA type of analysis that measures productivity in a field: counts such as annual publication numbers and the most productive author/institution/country.
Science mappingAn approach that displays the relationships among publications as visual networks (maps); it makes the structure and clusters of a field visible.
Systematic reviewA type of review that searches and evaluates the literature with a predefined, transparent, and reproducible method in order to answer a specific question. It provides deeper and more qualitative insight on small data sets.
Meta-analysisA method that statistically combines the numerical results of multiple studies; it sits between bibliometrics and the systematic review.
TriangulationCross-validating a finding with more than one method; for example, confirming a bibliometric result with a systematic review.
Gray literature / gray journalsPublications outside the traditional peer-reviewed channels (reports, small journals); the term “gray journals” is also used for some large publishers whose quality is debated.

B. Network and Co-occurrence Analyses

TermDefinition
Co-occurrence analysisA method that establishes a relationship by measuring how often two items (e.g., two keywords) appear together in the same documents.
Co-authorshipAn analysis that examines the collaboration network authors build through jointly authored articles.
Co-citation analysisMeasures how often two works are cited together by other works; those cited together are considered related.
Bibliographic couplingRelating two works by the degree to which they cite common sources; those with many shared references are considered similar.
ClusterA group formed by strongly connected items (publications, authors, keywords); it represents a sub-theme or research area.
ClusteringThe algorithmic process of dividing items into clusters based on their relationships; it is the basic building block of the map and must be interpreted by hand.
Link strengthA numerical value showing how strong the relationship between two items is.
Co-occurrence networkA visual network of nodes and links based on the co-occurrence relationships of terms.

C. Clustering and Visualization Parameters

TermDefinition
ThresholdThe minimum limit required for an item to be included in the analysis (e.g., a keyword appearing at least 5 times). It determines the size of the network.
Resolution parameterA value that adjusts how fine or coarse the clusters will be; high resolution produces many small clusters, low resolution a few large ones.
PruningThe process of removing weak links to simplify complex networks containing a large number of connections.
ModularityA metric measuring how well clusters are separated from one another; a high value means clearly separated clusters.
Silhouette valueA metric measuring the internal consistency (homogeneity) of a cluster; it shows how well items fit the cluster.
Betweenness centralityA metric that identifies key nodes acting as bridges in the network, connecting different clusters.
Citation burst / burst detectionDetecting moments when citations or terms increase sharply within a given period; it is a sign of rising trends.
ThesaurusA list that maps synonymous or differently spelled terms to one another; for example, it treats “AI” and “artificial intelligence” as the same concept.

D. Citation, Impact, and Counting Measures

TermDefinition
CitationA reference made by one work to another; it is the basic unit of measuring impact.
Citation analysisA method that measures impact and relationships by examining citation patterns.
NormalizationA correction applied to make data from different fields or periods comparable; it balances citation differences across fields.
h-indexA measure that combines an author’s productivity and impact in a single number: having h papers each cited at least h times.
Journal Impact Factor (JIF)A prestige indicator based on the average number of citations received by a journal’s articles over a given period.
Full countingA method of counting an article in full (as one) for each of its authors/institutions.
Fractional countingA method of counting an article’s contribution by dividing it among its authors/institutions (e.g., 1/3 to each of 3 authors).
Alternative metrics / altmetricsImpact indicators other than citations: downloads, shares, mentions in news/blogs, social-media engagement. They suffer from a lack of standardization and validation.
Dark citationsCitations in sources that do not enter traditional databases (e.g., policy documents); they are invisible in standard analyses.

E. Databases and Software Tools

TermDefinition
WoS (Web of Science)One of the two most common large citation databases; known for its selective indexing.
ScopusA large citation database that, together with WoS, dominates standard analyses.
Google Scholar (GS)A very broad academic search engine (including gray literature); however, bulk data extraction is difficult due to the lack of an API.
VOSviewerThe most widely used free tool for clustering and network visualization thanks to its intuitive interface; it lacks advanced metrics and time-series analysis.
CiteSpaceA powerful but hard-to-learn visualization tool that shows the development of fields over time and includes advanced metrics.
Bibliometrix / BiblioshinyA comprehensive R-based science-mapping tool (Biblioshiny is its visual interface).
SciVal / InCitesAuxiliary tools linked to Scopus and WoS respectively, providing basic performance and citation analyses and normalized comparisons.
API (Application Programming Interface)An interface that provides automatic access to a software’s data through programs; it is critical for bulk data downloading in bibliometrics.
ORCID / ResearcherID / ScopusIDIdentity systems that uniquely identify authors; they prevent name ambiguity.

F. Reporting Guidelines and Reform Initiatives

TermDefinition
PRISMAAn established, widely accepted reporting guideline for systematic reviews and meta-analyses. It serves as a model for the search for a similar standard in bibliometrics.
BIBLIOA reporting guideline proposed for bibliometric reviews of the biomedical literature (Montazeri et al., 2023).
GLOBALA 32-item general-purpose guideline for reporting bibliometric analyses (Ng et al., 2025).
PRIBAA preferred reporting items guideline for bibliometric analyses (used in health-medicine-focused evaluation).
VALORAn evaluation framework for readers and reviewers based on the principles of Validation, Alignment, Registration, Overview, and Reproducibility (Hoang, 2025).
DORAThe San Francisco Declaration on Research Assessment; it opposes citation-based metrics being the sole criterion in evaluation.
Leiden ManifestoA set of principles compiling best practices for the responsible use of bibliometric indicators (Hicks et al., 2015).
CoARAA coalition aiming to reform research assessment, advocating a shift from dependence on quantitative metrics to qualitatively grounded approaches.

G. Other Technical Terms

TermDefinition
PrecisionHow much of the search results are genuinely relevant to the topic; high precision means few “false” results.
RecallHow much of all the relevant articles are captured; high recall means few “missed” articles.
Research frontsThe most current and active topic clusters of a field; often determined by the term frequency of the citing articles.
Transient authorsResearchers, often from outside the field, who contribute only a single bibliometric article and then leave.
LDA (Latent Dirichlet Allocation)A topic-modeling technique that probabilistically extracts hidden themes in texts; used in qualitative theme discovery.
n-gramSequences of n consecutive words in a text; used in the numerical analysis of large bodies of text.
NMA (Network Meta-Analysis)An advanced type of meta-analysis that compares multiple treatments/options through indirect evidence.
LLM (Large Language Model)A large language model; AI systems that can automate data extraction and save time in analysis.
Search fields (ti, kw, ab, ts)Codes indicating a search in the title (ti), keyword (kw), abstract (ab), and broad topic (ts) fields, respectively.

Source Article and Principal Works Cited

Primary source: Ellegaard, O. (2026). The method of bibliometric analysis: a critical evaluation. Scientometrics. https://doi.org/10.1007/s11192-026-05670-6

The problems and recommendations in this document are compiled from the article above. The principal works on which the article bases its views are listed below (selected from the article’s reference list):

Ahmi, A. (2022). Bibliometric analysis for beginners. UUM Press.

Ansorge, L. (2024). Bibliometric studies as a publication strategy. Metrics, 1(1), 5.

Aria, M., & Cuccurullo, C. (2017). Bibliometrix: An R-tool for comprehensive science mapping analysis. Journal of Informetrics, 11(4), 959–975.

Ashley, J. (2025). Is the literature review paper dead? How AI is transforming the research landscape in DNA research. Nucleic Acid Insights, 3(1), 7–10.

Chen, C. (2006). CiteSpace II. Journal of the American Society for Information Science and Technology, 57(3), 359–377.

Chen, C. (2016). CiteSpace: A practical guide for mapping scientific literature. Nova Science Publishers.

Chen, H., Mehra, A., Tasselli, S., & Borgatti, S. P. (2022). Network dynamics and organizations. Journal of Management, 48(6), 1602–1660.

Cheng, K., et al. (2024). The rapid growth of bibliometric studies: A call for international guidelines. International Journal of Surgery, 110(4), 2446–2448.

Donthu, N., Kumar, S., Mukherjee, D., Pandey, N., & Lim, W. M. (2021). How to conduct a bibliometric analysis. Journal of Business Research, 133, 285–296.

Drew, C. H., et al. (2016). Automated research impact assessment. Scientometrics, 106, 987–1005.

Fei, X., & Hu, Y. (2024). A call for the establishment of bibliometric reporting guidelines. Journal of Clinical Nursing, 34, 3425–3426.

Glänzel, W. (1996). The need for standards in bibliometric research and technology. Scientometrics, 35(2), 167–176.

González-Alcaide, G. (2021). Bibliometric studies outside the information science and library science field. Scientometrics, 126(8), 6837–6870.

Haley, M. R. (2023). A dozen topics to debate before using bibliometrics to assess scholarly research articles. SSRN 4338192.

Harzing, A. W., & Alakangas, S. (2016). Google Scholar, Scopus and the Web of Science. Scientometrics, 106(2), 787–804.

Haustein, S., & Larivière, V. (2015). The use of bibliometrics for assessing research. In Incentives and performance (pp. 121–139). Springer.

Hicks, D., Wouters, P., Waltman, L., & de Rijcke, S. (2015). The Leiden Manifesto for research metrics. Nature, 520(7548), 429–431.

Hoang, A. D. (2025). Evaluating bibliometrics reviews: A practical guide for peer review and critical reading. Evaluation Review, 49(6), 1074–1102.

Hulland, J. (2024). Bibliometric reviews—some guidelines. Journal of the Academy of Marketing Science, 52(4), 935–938.

Jonkers, K., & Derrick, G. E. (2012). The bibliometric bandwagon. JASIST, 63(4), 829–836.

Keralis, J. M., Albertorio-Díaz, J., & Hoppe, T. (2023). Dark citations to Federal resources and their contribution to the public health literature. Frontiers in Research Metrics and Analytics, 8, 1235208.

Koo, M., & Lin, S. C. (2023). An analysis of reporting practices in the top 100 cited health and medicine-related bibliometric studies. Heliyon.

Kurutkan, M. N. (2026). GLOBAL: The first evidence-based reporting guideline for bibliometric analyses. Bibliometrics in Health Sciences.

Lazarides, M. K., Lazaridou, I. Z., & Papanas, N. (2025). Bibliometric analysis: Bridging informatics with science. The International Journal of Lower Extremity Wounds, 24(3), 515–517.

Li, J., Deacon, C., & Keezer, M. R. (2024). The performance of bibliometric analyses in the health sciences. Current Medical Research and Opinion, 40(1), 97–101.

Li, J., Deacon, C., & Keezer, M. R. (2025). Response to ‘Frontiers in bibliometric analysis…’. Current Medical Research and Opinion, 41(4), 669–670.

Lim, W. M., Kumar, S., & Donthu, N. (2024). How to combine and clean bibliometric data and use bibliometric tools synergistically. Journal of Business Research, 182, 114760.

Moed, H. F. (2020). Appropriate use of metrics in research assessment of autonomous academic institutions. Scholarly Assessment Reports, 2(1), 1.

Moher, D., et al. (2010). Preferred reporting items for systematic reviews and meta-analyses: The PRISMA statement. International Journal of Surgery, 8(5), 336–341.

Montazeri, A., et al. (2023). Preliminary guideline for reporting bibliometric reviews of the biomedical literature (BIBLIO). Systematic Reviews, 12(1), 239.

Montazerian, M., & Dorch, B. F. (2022). Quality and quantity in research assessment. Frontiers in Research Metrics and Analytics, 7, 991550.

Moral-Muñoz, J. A., et al. (2020). Software tools for conducting bibliometric analysis in science. El Profesional de la Información, 29(1), e290103.

Ng, J. Y., et al. (2025). Guidance for the reporting of bibliometric analyses: A scoping review (GLOBAL). Quantitative Science Studies, 6, 988–1001.

Page, M. J., et al. (2021). The PRISMA 2020 statement. BMJ.

Passas, I. (2024). Bibliometric analysis: The main steps. Encyclopedia, 4(2), 1014–1025.

Repiso, R., & Cabezas-Clavijo, Á. (2025). A critique of ‘quick and dirty’ bibliometrics. Revista Panamericana de Comunicaciones.

Synnestvedt, M. B., Chen, C., & Holmes, J. H. (2005). CiteSpace II: Visualization and knowledge discovery in bibliographic databases. AMIA Annual Symposium Proceedings, 2005, 724.

Tan, Y., et al. (2025). Angiogenesis after acute myocardial infarction: A bibliometric-based literature review. Frontiers in Cardiovascular Medicine, 12, 1426583.

Thompson, D. F., & Walker, C. K. (2015). A descriptive and historical review of bibliometrics with applications to medical sciences. Pharmacotherapy, 35(6), 551–559.

Torres-Salinas, D., Arroyo-Machado, W., & Robinson-García, N. (2025). Principles of evaluative bibliometrics in a DORA/CoARA context. InfluScience Ediciones.

van Eck, N. J., & Waltman, L. (2014). Visualizing bibliometric networks. In Measuring scholarly impact (pp. 285–320). Springer.

van Eck, N. J., & Waltman, L. (2023). VOSviewer manual (Version 1.6.20). Universiteit Leiden.

Wallin, J. A. (2005). Bibliometric methods: Pitfalls and possibilities. Basic and Clinical Pharmacology and Toxicology, 97(5), 261–275.

Wang, J. (2025). Frontiers in bibliometric analysis: The overrepresentation of bibliometric analyses in grey publisher journals. Current Medical Research and Opinion, 41(4), 667–668.

Watson, R., & Hayter, M. (2024). Bibliometric analyses: Do they contribute to knowledge? Journal of Clinical Nursing, 34, 1101–1102.

— End of document —

Appendices

1. On which points is there consensus?

The article is a critical text; but beneath the criticism lies common ground on which the field largely agrees. What is interesting is this: this consensus is broader than a “critics vs. practitioners” divide — even the developers of the tools converge on the same points.

Clear points of consensus:

  • Transparency and reproducibility are indispensable. The download date, search profile, software, and all parameter settings must be disclosed. The BIBLIO, GLOBAL, PRIBA, and VALOR guidelines all converge on this point. This is the strongest consensus in the field.
  • Citation count cannot be the sole measure of quality. This is the common denominator of DORA, the Leiden Manifesto, and CoARA; there is near-universal agreement at the institutional level.
  • Method choices materially change the result and must be justified. Among those who say so are VOSviewer’s own developers (van Eck & Waltman) and CiteSpace’s developer (Chen). That is, this is not an outside criticism but, in part, an acknowledgment from within.
  • A bibliometric study needs a genuine research question — the method itself cannot be the goal.
  • A common reporting standard is needed. The emergence of four separate guidelines in a short period shows not disagreement on the details but consensus on the need for a standard.
  • Bibliometrics is valuable not on its own but when combined with domain expertise and qualitative methods (triangulation).

So the consensus is deeper than it first appears; the debate is not about “is the method valuable?” but about “how is it legitimately applied?”

2. Which parts do editors problematize?

The editorial voice (Watson & Hayter 2024, Wang 2025, Li et al.) is markedly a gatekeeper perspective; it focuses on questions of value and venue rather than on technical parameter debates. The editors’ real discomfort is:

  • “Does this study belong in a peer-reviewed journal?” — This is the editors’ distinctive concern. The article’s most provocative suggestion (that narrow, descriptive analyses might be more suitable as reports/memos/institutional reports) comes directly from this editorial viewpoint.
  • The flow/volume problem — the pile arriving on their desks. That AI accelerates this flow is especially worrying.
  • The “colorful diagram illusion” — visual richness masking shallow content; descriptive studies offering no real contribution.
  • “Transient authors” — those who produce a single analysis and leave without engaging with the bibliometric literature.
  • The spread of “gray journals” (Frontiers, MDPI) and mass, low-budget production.
  • Strategic/career-oriented motives — “gaming” the citation return.

In short, editors foreground the problem of epistemic gatekeeping: “Is this knowledge, and does it belong here?” The deeper technical problems (threshold, resolution, pruning) are more the domain of the methodologists (van Eck & Waltman).

3. What should authors pay attention to?

This is the essence of the article’s minimum criteria for authors — it can be read as a kind of checklist:

  • Start with a clear research question; make the method serve the question.
  • Understand the metrics; do not run the software blindly. If you do not know how a number is calculated, do not report it.
  • Document everything: download date, search profile (precision/recall), database, all parameters, counting method (full/fractional), clustering and pruning decisions.
  • Be self-critical (reflexive): clearly state the method’s limits, sources of error, and statistical uncertainties.
  • Triangulate: corroborate findings based on large clusters with a systematic review or qualitative evaluation.
  • Consult a domain expert; where possible, include them as a co-author.
  • Produce contextualized insight rather than numbers and charts. The study must offer something new — updating old counts is not enough.
  • Normalize citations when comparing across fields.
  • Choose the right publication channel.

4. What should be done to make the method more useful?

These are systemic/structural improvements:

  • Convergence in reporting guidelines — the adoption and harmonization of standards such as GLOBAL.
  • Formal bibliometrics training — university curricula, CWTS (Leiden) courses, workshops. One of the root causes of the problems in this document is that the method is learned “self-taught and half-baked.”
  • Combining quantitative bibliometrics with qualitative/systematic methods — seeing bibliometrics not as a standalone product but as a component of an evidence-synthesis pipeline.
  • Context-sensitive citation tools (Scite AI, automated research-impact assessment) — a shift from the “quantity” of citation to its “quality.”
  • Author-identification infrastructure (ORCID, etc.) and better disambiguation.
  • Database interoperability — integration of field-specific sources (Medline, PsycInfo) and metadata standardization.
  • Cross-validation across tools — to see whether the result depends on the data or on the tool.

5. What should other stakeholders do?

The article focuses mainly on authors and editors; but when a full stakeholder ecosystem is considered, responsibility is distributed. This has a concrete counterpart because it is a multi-role field that you, too, are part of:

  • Reviewers: use reader/reviewer-focused frameworks such as VALOR; treat transparency and reproducibility as conditions; involve at least one domain expert in the evaluation. (When writing a reviewer report, these frameworks provide a concrete checklist.)
  • Database providers (Clarivate/WoS, Elsevier/Scopus): metadata standardization, better disambiguation, API access — the solution to Google Scholar’s “rich but unusable” problem is in their hands.
  • Software developers (VOSviewer, CiteSpace): better defaults, documentation, and making parameter effects transparent.
  • Universities and institutions: provide formal training; do not over-rely on metrics in hiring/promotion. (For someone conducting academic appointment/promotion evaluation, the DORA/CoARA principles directly enter the criteria set.)
  • Research-assessment bodies and funders: participation in DORA/CoARA, diversified and responsible use of metrics (Leiden).
  • Journals and publishers: develop editorial policies on where a bibliometric study belongs.
  • The bibliometrics community: build standards and networks, bridging those who know the method (informetrics) with those who know the subject (medicine, health management).
  • Policymakers: as bibliometrics increasingly feeds policy briefs, they too need bibliometric literacy.

6. Additional Considerations

Beyond what the article says, a few patterns stand out:

  • The reflexivity paradox. A text criticizing bibliometrics is itself a kind of meta-literature analysis; it uses the field’s tools against the field. Moreover, the article demonstrates its own thesis on itself by examining the same 2023 data with two different methods and producing different themes. This is a rhetorically powerful but also epistemologically uncomfortable loop.
  • The “quantitative theater” problem. The deepest point, in my view, is this: the appearance of objectivity given by numbers and maps conceals the underlying subjective choices (threshold, resolution, thesaurus). This is in fact a measurement-validity problem — the construct validity of a bibliometric “map” should be questioned, just as for a latent construct. For someone coming from the SEM/scale-development tradition, the analogy is quite direct: parameter selection is a kind of “model specification” decision, and when it is not reported, the result cannot be validated.
  • A single root cause: the incentive structure. Most of the technical problems (wrong parameters, shallow production, gray journals) are in fact symptoms. The root cause is single: publish-or-perish pressure + high citations + low cost + the ease of AI. That is why real reform is not methodological but institutional (DORA/CoARA). Writing guidelines is necessary but not sufficient; unless the evaluation culture changes, the problem persists.
  • The AI inflection point. AI lowers the method threshold even further, and this cuts both ways: it will either flood journals with shallow analyses or solve some problems (context-sensitive citation, automated screening, disambiguation). The net effect depends on governance — which intersects directly with AI-governance debates.
  • The standardization-race risk. The very “standard” proposed as the solution to fragmentation is itself fragmenting: there are now four rival guidelines (BIBLIO, GLOBAL, PRIBA, VALOR). The next need is convergence and harmonization — and as the author of GLOBAL, you are already at that table.
  • Health sciences as the epicenter. The criticism is especially sharp for the health sciences (gray journals, the volume from China, “dark citations” in clinical guidelines). This is also an opportunity: health management/policy is precisely the field where bibliometrics can generate real value as decision support rather than description — question-driven analyses that feed institutional reports and policy briefs. The article’s suggestion to “disseminate through the right channel” can turn into a directly applicable strategy for your publication ecosystem.

Incorrect or debatable claims

1. The misuse of the “gray journal” concept — and a hidden internal inconsistency. The article labels Frontiers and MDPI as “gray journals.” This is a terminological slippage: by its established definition, grey literature refers to non-commercial sources that do not go through peer review (reports, theses, working notes). Frontiers and MDPI, by contrast, are indexed, peer-reviewed, commercial journals. More importantly, the inference “publisher → quality” is the very proxy reasoning the article condemns elsewhere. A text that says “citations ≠ quality” implicitly assumes “publisher = quality.” This is the same problem it tries to solve; moreover, the “especially of Chinese origin” framing rests on empirically and ethically fragile ground by equating volume and geographic origin with quality.

2. The “transient authors” judgment contradicts its own advice. On the one hand, the article wants domain expertise included in bibliometrics; on the other, it disparages “transient authors” who produce a single analysis and leave. Yet these two groups are largely the same: a cardiologist analyzing their own field once can often produce a more valid study than a bibliometrics expert who does not know the field. The “transience” stigma takes embeddedness in bibliometrics as the measure of quality; but the article also praises embeddedness in the subject. This tension is left unresolved.

3. The assumption “speed + cheapness + volume = low quality” is value-laden, not empirical. This triad is used almost as an axiom throughout the article; yet the correlation is not proven. Being cheap and fast also democratizes the method — it grants access to researchers in resource-limited settings (the Global South). This stance, which equates accessibility with corruption, has a mildly elitist/gatekeeping tone and is presented as if it were empirical.

4. “Descriptive analysis is short-lived, therefore low-value” is a logical leap. A snapshot can be valuable precisely because it is dated: repeated snapshots produce scientifically valuable time series. Epidemiological surveillance or census data “ages” but is foundational. The article underestimates the cumulative value of repeated descriptive measurement.

5. The “actually qualitative” overcorrection. That parameter choices are subjective is true; but jumping from there to “bibliometrics is not objective counting, it is actually qualitative” is too much. Every quantitative method has researcher degrees of freedom — a regression also has specification decisions. The honest framing should have been “bibliometrics, like every statistical method, involves modeling decisions,” not “secretly qualitative.” This is an overcorrection.

6. The “send it to a report/memo” advice ignores the very incentive reality it diagnoses. The article identifies “publish or perish” pressure as a problem; but its solution is to ask the individual to turn toward a channel that this system penalizes. In many academic systems, including Turkey’s, only peer-reviewed journal output “counts.” Advising the individual to bear the system’s cost alone — without changing the system — is somewhat naive.

What the article leaves out

1. The absence of empirical evidence for its own claims (the greatest irony). This text, which criticizes a quantitative method, is largely opinion/narrative in nature. “Most studies lack a research question,” “a large portion do not reach maturity,” “gray journals harbor them” — but how many? What proportion? There is no content analysis, sample, or percentage. The article partly commits the same sin of “claims that are not methodologically rigorous.” A meta-research/empirical audit would have strengthened its thesis many times over.

2. Open-science infrastructure and open citation data. A call for reproducibility is made; but the concrete solution — uploading the query, the data set, the thesaurus file, and the R/Biblioshiny script to a repository such as Zenodo/OSF — is barely addressed. More strikingly: OpenAlex and OpenCitations are not mentioned at all. Yet OpenAlex directly solves the Google Scholar “no API” problem the article complains about; open citation data is also the antidote to the coverage and reproducibility complaints. Omitting these is a real gap — and in the design of a reporting guideline (e.g., the next version of GLOBAL) this is exactly the place to fill: not reporting parameters in prose, but sharing data and code.

3. Uncertainty and statistical stability. Bibliometric counts are treated as if fixed; yet they carry sampling and measurement error. The article touches on “statistical uncertainty” in passing but does not develop it: there is no notion that confidence intervals, bootstrap, sensitivity, and robustness analyses are needed for the indicators. Reporting how robust a map is to parameter changes — the same logic as the bootstrap validation in your health-economics studies — is a fundamental omission.

4. Structural bias and performativity. The article addresses coverage and multiple indexing; but it does not treat systematic biases (English dominance, the Matthew effect in citations, gender/geography inequalities in authorship data) and how these become embedded in the “maps” and reproduced as if objective. Bibliometrics can launder structural inequalities under the guise of neutral “evidence.” A second, related gap is performativity (Goodhart’s/Campbell’s law): once an indicator becomes the target in evaluation, it ceases to be a measure. The article touches on DORA but does not theorize how bibliometric analysis itself feeds the incentive structure it criticizes, as a self-reinforcing loop.

5. AI’s positive methodological role and its governance. AI is framed almost solely as a “threat that lowers the threshold.” Its positive methodological roles are missing: LLM-assisted screening, semantic disambiguation, citation-context classification, automated PRISMA, and embedding-based topic models that go beyond keyword co-occurrence. Moreover, AI governance for bibliometrics (provenance, hallucinated references, the reproducibility of AI-assisted pipelines) is not addressed at all — an open future agenda that intersects directly with your interest in AI governance.

6. Decision orientation and the absence of positive examples. The article is almost entirely diagnostic; it offers no decision-theoretic framework of “useful for whom, for which decision?” Yet such a framework would have sharpened the “research question” critique (and aligns with the information-value thinking in health management). It also does not present well-executed studies that could serve as models; without a model to emulate, only criticism remains.

7. New methods and field-specific norms. VOSviewer/CiteSpace (co-citation, co-word) are placed at the center; network embeddings, dynamic topic models, citation-context/sentiment analysis, and full-text mining are largely left out. The prescription is also too general: how bibliometric norms in nursing, physics, and business differ (citation density, database coverage, conventions) is not developed.

In short: the article points to the right questions but does so partly in the very style it criticizes — unevidenced, overly general, and value-laden. The most productive reading is to treat it not as the last word but as a starting text that someone like you — who writes guidelines and conducts evaluations — would empirically test and complete along the axes of open science, uncertainty analysis, and AI governance.

Subscribe to the Health Topics Newsletter!

Google reCaptcha: Invalid site key.