Scopus vs Web of Science vs Google Scholar: How Indexing Differs (2026)
Three databases dominate academic literature discovery and research metrics — yet a researcher submitting a grant application, a PhD candidate building a systematic review, and a librarian advising on journal quality assessments can arrive at radically different conclusions depending on which platform they consult. The divergence is not a measurement error. It reflects structural decisions each platform has made about what constitutes scholarly literature, who is authorised to enter the index, and how citations are counted. Understanding Scopus vs Web of Science indexing — and how Google Scholar fits between and beyond them — is therefore a prerequisite for rigorous academic work in 2026.
This article provides a research-grade comparison of all three databases, drawing on official content-policy documentation from Elsevier and Clarivate as well as peer-reviewed bibliometric studies, most notably the landmark analysis by Martín-Martín, Orduna-Malea, Thelwall, and López-Cózar published in the Journal of Informetrics. The goal is not to crown a winner but to give researchers a sufficiently detailed map of each index so they can deploy the right tool for each task.
The Fundamental Architecture of Each Index
Before comparing coverage statistics, it helps to understand what each platform is architecturally. Scopus, owned by Elsevier and launched in 2004, is a curated abstract and citation database. Every record in Scopus has been evaluated by the Content Selection and Advisory Board (CSAB) before the source journal is admitted. Scopus does not index the full text; it indexes bibliographic metadata and abstracts, with citation links derived from reference lists in the indexed material.
Web of Science, originally the Science Citation Index created by Eugene Garfield at the Institute for Scientific Information and now operated by Clarivate, follows a comparable curation model but applies a different — and by its own account more stringent — editorial process. Its premium tiers (Science Citation Index Expanded, Social Sciences Citation Index, Arts and Humanities Citation Index) are restricted to journals that demonstrate not only editorial quality but measurable citation impact. Journals meeting quality but not impact thresholds enter the Emerging Sources Citation Index (ESCI), a distinct tier introduced in 2015.
Google Scholar is structurally different. It is a web crawler, not an editorially curated database. Google’s bots discover documents that appear to be scholarly — journal articles, conference papers, preprints, theses, dissertations, book chapters, and technical reports — and index them automatically. There is no CSAB equivalent, no impact evaluation, and no formal acceptance process. This architecture makes Google Scholar the broadest of the three but also the least consistent in metadata quality.
Selection Criteria: How Journals Enter Each Database
Scopus: The CSAB Process
According to Elsevier’s official content-policy documentation, journals submitted for Scopus review must meet a set of technical prerequisites before reaching the advisory board: peer-reviewed content with a publicly described review process, regular publication, a registered ISSN, and full-text accessibility by Scopus editors. Thousands of titles are suggested each year; Elsevier reports that approximately one-third meet the technical criteria, and of those roughly half are ultimately accepted after CSAB review. As of August 2024, Elsevier removed the previous two-year minimum publication period, allowing newer journals to submit earlier — though publishers are advised to calibrate submission timing against the volume of content already published.
The CSAB evaluates titles against both quantitative criteria (citation activity of published articles) and qualitative criteria (editorial board composition, editorial scope clarity, author and citation geographic diversity, and absence of problematic publication practices). Retracted and de-listed titles are periodically removed from the active index.
Web of Science: 28 Criteria, Four Impact Thresholds
Clarivate’s evaluation framework for the Web of Science Core Collection specifies 28 criteria: 24 quality criteria addressing editorial rigor and best practices, and four impact criteria assessing a journal’s citation influence within its field. The quality criteria include requirements for transparent peer review, full-text access for the editorial team, consistent publication scheduling, ethical guidelines conformance, and the absence of self-citation manipulation. Impact criteria, while currently stated as secondary to quality for initial evaluation, determine which tier — SCIE, SSCI, AHCI, or ESCI — a journal occupies.
This tiered structure is consequential for researchers. A journal’s inclusion in ESCI means it appears in Web of Science search results but is not counted in Journal Citation Reports (JCR) Impact Factors. Many institutional research assessments, including the UK Research Excellence Framework’s journal lists and analogous national evaluation exercises, distinguish between SCIE/SSCI journals and ESCI-only journals. Researchers targeting journals for publication with formal assessment implications need to verify the specific tier, not merely Web of Science presence.
Google Scholar: No Submission Process
Google Scholar has no submission process. A document becomes indexed when Google’s crawler encounters it in a location that signals scholarly provenance — a publisher’s website, an institutional repository, an author’s personal academic page, or a preprint server such as arXiv. Publishers can improve discoverability by following Google Scholar’s indexing guidelines (correct meta-tags, stable URLs, structured bibliographic metadata), but cannot guarantee or reject indexing. This approach produces exceptional breadth at the cost of quality control: duplicate records for the same paper at different repository locations, citations to retracted papers, and non-peer-reviewed documents treated as equivalent to journal articles all coexist in the index.
Coverage Scope: Journals, Disciplines, Languages, and Document Types
Raw Journal Numbers
| Database | Active journals (approx.) | Total documents (approx.) | Selection model |
|---|---|---|---|
| Scopus | ~27,000 | ~90 million | Editorial curation (CSAB) |
| Web of Science Core Collection | ~21,000 | ~75 million | Editorial curation + impact tiers |
| Google Scholar | Not disclosed | ~160 million+ | Automated web crawl |
These headline numbers deserve two caveats. First, Scopus and Web of Science figures change as journals are added and de-listed; verify current counts against official coverage guides when precision matters for a methods section. Second, “documents” in Google Scholar includes document types that Scopus and Web of Science would not count as scholarly publications — working papers, technical reports, thesis chapters, and blog posts that happen to carry bibliographic metadata — so the 160-million figure reflects a fundamentally different definition of the index unit.
Geographic and Language Coverage
Both Scopus and Web of Science have historically skewed toward journals published in English and by large Western commercial publishers. This is a well-documented bibliometric concern. Scopus has actively pursued international and regional titles in recent years, which partly explains its higher raw journal count relative to Web of Science. A 2024 preprint published on arXiv examining variations in journal coverage between 2001 and 2020 found that both databases expanded coverage substantially over the period, with Scopus showing faster growth in non-English titles.
Google Scholar’s geographic coverage is structurally broader because it harvests documents from regional institutional repositories, national open-access portals, and non-commercial academic websites that would not meet the submission and quality thresholds of either curated database. For researchers working in fields with strong regional publication traditions — legal scholarship, education research in non-OECD contexts, regional history — Google Scholar may be the only automated tool that captures a representative slice of the literature.
Document Types
All three databases index journal articles as their primary document type. Beyond that, coverage diverges:
- Conference proceedings: Scopus indexes conference proceedings extensively, particularly for engineering, computing, and life sciences through publishers such as IEEE and ACM. Web of Science includes selected conference proceedings through its CPCI-S and CPCI-SSH indexes. Google Scholar indexes conference papers from nearly any institution that publishes PDFs online, making it the dominant source for computer-science conference citations.
- Books and book chapters: Scopus indexes book series and selected book chapters. Web of Science expanded its Book Citation Index in 2011. Google Scholar has the most comprehensive book coverage, though metadata inconsistencies are common.
- Preprints: Scopus and Web of Science have begun indexing preprints from arXiv, bioRxiv, medRxiv, and SSRN, but coverage is incomplete and preprints are flagged rather than treated as equivalent to peer-reviewed publications. Google Scholar indexes preprints without such differentiation, which inflates citation counts and requires researchers to verify the publication status of highly cited papers.
- Theses and dissertations: Google Scholar indexes theses from institutional repositories; Scopus and Web of Science do not index theses as a document type.
- Grey literature: Google Scholar captures a substantial volume of grey literature; the curated databases do not systematically include it.
Why Citation Counts Diverge
The single most confusing aspect of working with multiple databases is that the same paper will carry different citation counts in each. The explanation is straightforward once the architecture is understood: each platform counts only citations that originate from within its own source pool.
If Paper A, published in a Scopus-indexed journal, is cited by Paper B in a journal indexed in Scopus but not in Web of Science, Scopus records that citation but Web of Science does not. If Paper A is also cited in a thesis hosted on a university repository, neither curated database records it — but Google Scholar does.
The landmark empirical treatment of this question is the Martín-Martín, Orduna-Malea, Thelwall, and López-Cózar (2018) study in the Journal of Informetrics, which systematically compared citations across 252 subject categories using approximately 2.5 million citation records. The study found that Google Scholar citation data is effectively a superset of both Scopus and Web of Science: almost all citations found in the curated databases also appear in Google Scholar, but Google Scholar adds substantial additional coverage not present in either. Spearman rank correlations between Google Scholar and Web of Science citation counts were generally high (ranging from 0.78 to 0.99), indicating that while absolute counts differ, the relative ranking of papers by citation impact is broadly consistent. The correlations were lower in the humanities, where Google Scholar’s broader coverage of books and grey literature creates the most divergence from the curated databases.
For practical purposes, this means:
- A paper’s Web of Science citation count will almost always be the lowest of the three.
- Google Scholar citation counts will almost always be the highest.
- The gap between databases is smallest in natural sciences and largest in humanities, social sciences, and interdisciplinary fields.
- Absolute citation numbers from different databases should never be compared directly without reporting which database was the source.
H-Index Variation Across Databases
Because the h-index is computed from a researcher’s citation record, and citation records differ across databases, the h-index is database-specific. A researcher’s Scopus h-index, Web of Science h-index, and Google Scholar h-index are three different measurements — not three attempts to measure the same thing.
The direction of variation is consistent: Google Scholar h-index values are typically highest, Web of Science values typically lowest, and Scopus in between. The magnitude of the gap varies considerably by discipline, career stage, and the extent to which a researcher’s work appears in document types that only Google Scholar indexes comprehensively (conference proceedings, preprints, books, theses). The article What Is the h-Index and How Do You Calculate It? on this site covers the calculation methodology and cross-database benchmarks in detail.
From a practical standpoint, researchers should:
- Always specify which database they used when reporting their h-index in a CV, grant application, or promotion document.
- Maintain a Google Scholar profile to capture the broadest citation record, but use Scopus or Web of Science for formal institutional reporting where curated data is required.
- Be aware that institutional research-assessment exercises (Research Excellence Framework in the UK, Excellence in Research for Australia, similar exercises elsewhere) typically mandate the use of curated-database metrics, not Google Scholar values.
Disciplinary Depth and Blind Spots
Natural Sciences and Medicine
Both Scopus and Web of Science provide strong coverage of the natural sciences, medicine, and engineering. Web of Science’s Science Citation Index Expanded is the historical backbone of biomedical citation analysis and is the source of the Journal Impact Factor computed by Clarivate’s Journal Citation Reports. PubMed/MEDLINE, the National Library of Medicine’s specialised biomedical index, complements both platforms and is the primary source for clinical systematic reviews; Scopus’s coverage of Embase (also Elsevier) provides additional pharmaceutical and clinical literature. Researchers in these fields who rely solely on Google Scholar risk including retracted papers that have been removed from curated databases.
Social Sciences
Coverage of social-science journals is reasonably strong in both curated databases. Web of Science’s Social Sciences Citation Index covers established journals; Scopus’s broader journal intake means it often captures newer social-science titles earlier. The most notable gap in both is a relative under-representation of research published in non-English-speaking countries and in government or policy reports that would be classified as grey literature — material that Google Scholar captures more fully. For research areas with significant policy relevance (public health, education, social welfare), supplementing database searches with targeted grey-literature searches is standard best practice. The article What Is Grey Literature in Research? on this site addresses this complementary source type in detail.
Humanities
The humanities present the clearest case for departing from curated-database primacy. Web of Science’s Arts and Humanities Citation Index covers selected humanities journals, but the humanities publish more substantially in books and book series — document types for which all three databases provide limited structural coverage. Google Scholar’s broad capture of books and book chapters makes it proportionally more useful for humanities researchers than for natural scientists. Citation analysis based on Web of Science or Scopus data alone systematically underestimates the scholarly impact of humanities researchers compared with their STEM counterparts, a concern that has driven initiatives such as the European Reference Index for the Humanities (ERIH PLUS) as a discipline-appropriate supplementary source.
Computer Science and Engineering
In computer science, the most prestigious publication venues are often conferences, not journals. The ACM Digital Library and IEEE Xplore are the canonical archives. Both Scopus and Web of Science index proceedings from these publishers, but Google Scholar’s coverage of computer-science conference papers is significantly more comprehensive and faster. For h-index computation in computer science, Google Scholar values are typically much higher than curated-database values because the bulk of citations come from conference-to-conference citation chains that Google Scholar captures and curated databases do not fully track.
Which Database to Use for a Literature Review
The consensus recommendation in systematic review methodology, reflected in the PRISMA 2020 guidelines and reinforced by database coverage studies, is that no single database provides sufficient coverage for a rigorous systematic literature review. The appropriate combination depends on the review’s disciplinary focus. When choosing AI-assisted tools to complement database searching, the comparison of Elicit vs Consensus vs Scite for evidence-synthesis AI provides an up-to-date assessment of which platforms best complement Scopus, Web of Science, and Google Scholar searches in 2026.
| Discipline / Review type | Primary databases | Supplementary sources |
|---|---|---|
| Clinical / biomedical | PubMed/MEDLINE, Embase (Scopus) | Cochrane Library, Web of Science, Google Scholar |
| Social sciences / psychology | Scopus, Web of Science (SSCI), PsycINFO | Google Scholar, grey literature databases |
| Humanities | Google Scholar, JSTOR, ERIH PLUS journals | Web of Science (AHCI), Scopus, institutional repositories |
| Computer science / engineering | ACM Digital Library, IEEE Xplore, Scopus | Google Scholar, Web of Science (CPCI-S) |
| Interdisciplinary / broad | Scopus + Web of Science | Google Scholar, domain-specific databases |
Constructing effective search queries for a multi-database review requires adapting your Boolean logic to each platform’s syntax and controlled-vocabulary system. The guide to Boolean search operators in academic databases on this site provides platform-specific examples for Scopus, Web of Science, PubMed, and Google Scholar. A fuller treatment of the multi-database literature review workflow appears in the systematic literature review guide.
One often-overlooked strategy in multi-database searching is citation chasing: using the reference lists of known relevant papers (backward chasing) and tracking all papers that have subsequently cited them (forward chasing). All three databases support this, but with different affordances. Scopus and Web of Science have structured “cited by” functionality that allows cited-reference searching as a search type. Google Scholar’s “cited by” link is useful but less structured. The comparison of citation chasing tools on this site assesses these and specialist tools such as Connected Papers and citationchaser.
Which Database to Use for Research Assessment and Metrics
For formal research-assessment purposes, the choice between Scopus and Web of Science is typically dictated by institutional context rather than researcher preference.
Clarivate’s Journal Citation Reports, which publishes the Journal Impact Factor and the CiteScore-equivalent Eigenfactor, draws exclusively from Web of Science Core Collection data. Impact Factors are computed only for journals in SCIE, SSCI, and AHCI — not ESCI. If a researcher’s institution uses Impact Factor as a journal-quality proxy, Web of Science indexing status and tier are the operative variables.
Scopus publishes its own journal-level metric, CiteScore, computed annually. CiteScore uses a four-year citation window (compared with the two-year window of the traditional Impact Factor) and counts citations from all Scopus-indexed sources, which generally yields higher raw values. Scopus also produces the Source Normalized Impact per Paper (SNIP) metric, which adjusts for citation density differences between fields — a useful tool for cross-disciplinary comparisons.
Neither Google Scholar nor its publisher-provided derivatives (like h5-index in Google Scholar Metrics) are typically accepted for formal institutional reporting, although researcher-level Google Scholar profiles are widely used informally and are particularly important in fields such as computer science and humanities where Google Scholar provides substantially better coverage than the curated databases.
Researchers preparing for promotion, grant applications, or national research assessments should maintain profiles in all three systems but ensure they can produce curated-database reports on demand. Scopus Author Profiles and Web of Science ResearcherID/ORCID integration allow researchers to claim and verify their publication records; claiming these profiles early in a research career and correcting name disambiguation errors prevents systematic under-counting in curated databases.
A Practical Multi-Database Workflow
Given the structural differences outlined above, researchers benefit from a tiered workflow rather than database loyalty. The following framework is applicable across disciplines:
Step 1: Define your scope and identify your primary databases
Consult your discipline’s systematic review reporting norms (PRISMA for health, ROSES for environmental sciences, ENTREQ for qualitative synthesis) to determine which databases are expected. Note that reviewers and examiners familiar with your field will check your methods section against these norms.
Step 2: Construct and test your search strategy in the most sophisticated interface
Scopus’s Advanced Search and Web of Science’s Advanced Search both support field-tagged queries, proximity operators, truncation, and controlled vocabulary (MeSH for biomedical topics in PubMed; the Web of Science subject heading system). Build your string in the platform with the strongest interface, then adapt it — not copy-paste it — to other platforms, accounting for syntax differences. Reviewing the expert techniques for Google Scholar advanced search will help you adapt your query for Google Scholar’s more limited but still powerful operator set.
Step 3: Run in all primary databases and de-duplicate
Export references from each database in a format compatible with your reference manager (RIS or BibTeX for Zotero and Mendeley; RIS for EndNote). Deduplicate systematically before screening — Zotero’s “Find Duplicates” function is a useful starting point, though manual review is always required because the same paper may appear under slightly different titles or with different author-name rendering across databases.
Step 4: Run Google Scholar as a supplementary sweep
Because of the practical limitations of Google Scholar — no bulk export, maximum 1,000 results per query, no structured deduplication — it works best as a validation tool rather than a primary search source. Run your key terms and check whether the top results include any sources not captured by your primary database searches. Any relevant additions should be checked for journal indexing status to understand why they were missed.
Step 5: Conduct citation chasing on key papers
For papers that turn out to be central to your review, run both backward citation chasing (scanning reference lists) and forward citation chasing (following “cited by” links in Scopus, Web of Science, and Google Scholar) to identify any relevant literature that your keyword strategy missed. This is particularly important in rapidly evolving fields where recent papers may not yet be well-indexed.
Step 6: Document your database choices in your methods section
Specify which databases you searched, the dates of your searches, and any date or language restrictions applied. PRISMA 2020 requires this at a minimum. Reviewers in health and social sciences increasingly also expect you to provide your full search string as supplementary material.
Key coverage figures (2026): Scopus ~27,000 journals / ~90M documents; Web of Science Core Collection ~21,000 journals / ~75M documents; Google Scholar ~160M+ documents (all types). Approximately 80–85% of WoS journals are also in Scopus; only ~60–65% of Scopus journals are in WoS. Source: Geographical and disciplinary coverage of OA journals: OpenAlex, Scopus, and WoS (PLOS ONE, 2025).
Frequently Asked Questions
Does Scopus or Web of Science index more journals?
As of 2026, Scopus indexes approximately 27,000 active peer-reviewed journals compared with Web of Science Core Collection’s roughly 21,000. Scopus therefore has broader raw journal coverage, particularly across non-English-language and regional titles. However, Web of Science applies stricter editorial selection criteria — requiring citation impact evidence for its premium tiers (SCIE, SSCI, AHCI) — so its Core Collection is widely considered more selectively curated. For research assessment exercises that reference Web of Science tiers (SCIE, SSCI), a journal’s absence from the higher Web of Science tiers carries specific implications that Scopus presence alone does not resolve.
Why are my citation counts different in Scopus, Web of Science, and Google Scholar?
Each database counts only citations from sources within its own index. Scopus tallies citations from its ~27,000 journals; Web of Science from its ~21,000; Google Scholar from a far broader pool including preprints, theses, and grey literature. Because the source pools differ, a paper in a journal not indexed in Web of Science will receive no Web of Science citation credit, yet its citations may appear in Scopus and almost certainly in Google Scholar. The Martín-Martín et al. (2018) study confirmed that Google Scholar citation counts are essentially a superset of the other two, particularly in the humanities and social sciences.
Which database should I use for a systematic literature review?
Major systematic review reporting guidelines, including PRISMA 2020, recommend searching multiple databases rather than relying on any single source. For clinical and life-science reviews, PubMed/MEDLINE plus Embase (accessible via Scopus) forms the standard pair. Scopus and Web of Science together cover most natural-science and social-science literature. Google Scholar is recommended as a supplementary source for grey literature and to catch citations missed by the curated databases. For humanities or regional-language topics, Google Scholar becomes proportionally more important.
Is Google Scholar reliable for research metrics?
Google Scholar provides broad coverage and is free, but it is less reliable for formal research metrics because its automated indexing can include duplicate records, preprint versions that inflate citation counts, and sources that would not pass editorial scrutiny in Scopus or Web of Science. For grant applications or academic promotion cases where formal bibliometric evidence is required, Scopus or Web of Science are the standard platforms. Google Scholar profiles are nonetheless useful for tracking all publicly available citations of your work, especially in fields such as computer science and humanities where curated-database coverage is weakest.
Does my h-index differ between Scopus, Web of Science, and Google Scholar?
Yes, often substantially. Because each database counts citations only from its own indexed sources, the h-index computed from Google Scholar is typically the highest, followed by Scopus, then Web of Science. The gap is most pronounced in fields with significant conference proceedings (computer science, engineering) and in the humanities, where Google Scholar captures citations from books, theses, and grey literature that the curated databases do not index. The article on this site covering the h-index in detail explains the calculation and these cross-database differences.
Does Web of Science cover conference proceedings?
Yes. Web of Science Core Collection includes the Conference Proceedings Citation Index — Science (CPCI-S) and Conference Proceedings Citation Index — Social Science & Humanities (CPCI-SSH), which index selected conference proceedings subject to the same editorial review as journals. Scopus also indexes conference proceedings, particularly for computer science and engineering through its inclusion of IEEE, ACM, and similar publishers. Google Scholar indexes conference papers extensively, making it especially valuable for computer science researchers where top venues are conferences rather than journals.
Write your thesis with AI
Structure, draft, cite, and format your thesis faster with Tesify’s AI writing tools, automatic bibliography, and plagiarism checker. Free to start, no credit card required.






Leave a Reply