How to Run a Bibliometric Analysis: A Step-by-Step Guide
A bibliometric analysis turns thousands of papers into a map of a field. This guide walks the full pipeline — and the decisions that determine whether the map means anything.
What bibliometric analysis is — and is not
Bibliometrics applies quantitative methods to publications and their citations to describe a research field. It answers questions a narrative review cannot: which authors, journals and countries dominate; how the literature has grown; which papers form the intellectual base; and how themes cluster and evolve. It is not a substitute for reading — it tells you the shape of a forest, not the content of individual trees. The best reviews combine both.
Step 1 — Define the question and scope
Decide what you want the map to show: the structure of a field, its evolution over time, an emerging theme, or a comparison of sub-areas. Your question dictates everything downstream. Pin down the time window, the document types (articles and reviews, usually), and the languages you will include.
Step 2 — Design the search query
This is the step that makes or breaks the study. Build a Boolean query from your core concepts and their synonyms, search title-abstract-keywords, and iterate. Too narrow and you miss the field; too broad and you drown in noise. Document the exact query string and the date you ran it — reviewers will ask, and reproducibility demands it.
Step 3 — Export the records
Export full records with cited references from Scopus or Web of Science (the citation data is what enables co-citation analysis). Biblio Infinity also reads PubMed (.nbib), BibTeX and RIS. If you search more than one database, export each separately so you can de-duplicate properly.
Step 4 — Clean and de-duplicate
Raw exports are messy: the same author appears as "Smith, J." and "Smith, John"; the same paper appears in two databases. Merge author and source name variants, remove duplicates, and screen out off-topic records. If you are doing a systematic review, document this screening as a PRISMA flow (records identified → duplicates removed → screened → included). Clean data is the difference between a map and a mess.
Step 5 — Performance analysis
Describe the field's productivity and impact:
- Production over time — the annual publication count and growth rate.
- Most productive and most cited authors, journals, institutions and countries.
- Impact indices — total and average citations, and the h- and g-index for authors or sources.
- Bradford's law to find the core journals, and Lotka's law to describe author productivity distribution.
These numbers must be computed from the actual records, not estimated — small counting errors propagate into wrong conclusions, which is why Biblio Infinity computes them in code rather than by approximation.
Step 6 — Science mapping
Mapping reveals structure and relationships:
- Co-citation analysis — papers frequently cited together form the intellectual foundation of the field.
- Co-word / keyword co-occurrence — terms that appear together reveal the conceptual structure and let you cluster themes.
- Co-authorship — the collaboration network among authors, institutions or countries.
- Thematic (Callon) map — plots themes by centrality and density into motor, niche, emerging/declining, and basic themes. This is the single most useful figure for spotting research gaps.
Step 7 — Interpret and write
A network diagram is not a finding; your interpretation is. For each map, write what it means: which cluster is the established core, which is emerging, where two clusters fail to connect (a gap and an opportunity). Tie the quantitative structure back to your reading of the key papers, and end with a concrete future-research agenda. That synthesis is the publishable contribution.
Common pitfalls
- Skipping de-duplication across databases, which double-counts everything.
- Over-interpreting tiny clusters built from a handful of papers.
- Reporting networks without interpretation — pretty pictures, no argument.
- Not reporting the exact query and date, making the study irreproducible.
Final thought
A bibliometric analysis is only as good as its query and its cleaning. Get those right, compute the metrics accurately, and the maps will tell a story worth publishing. Biblio Infinity handles the computation so you can spend your time on the part that matters — interpretation.
Run the analysis on your export
Drop a Scopus, Web of Science, PubMed, BibTeX or RIS file into Biblio Infinity and it computes performance metrics and science maps in your browser — accurate, reproducible, and free.