Before you conduct a single interview or design a survey, the smartest move in any research project is to find out what already exists. In communication research, this means mastering secondary data – information that has already been collected, processed, and published by someone else. According to the SAGE Encyclopedia of Communication Research Methods, secondary data refers to data collected by someone other than the researcher, originally gathered for a different purpose but highly useful for new inquiries. For students and professionals in Journalism and Mass Communication, knowing exactly where to find this data – and how to evaluate it – is one of the most practical skills you can develop.
Table of Contents
- What counts as secondary data?
- Why communication researchers must cast a wide net
- The internet as a secondary data source
- Using search operators effectively
- The Internet Archive and digital history
- Academic databases: the deep ocean of secondary data
- Google Scholar
- JSTOR
- ERIC
- Communication & Mass Media Complete (CMMC)
- Scopus and Web of Science
- Official publications and government data
- Libraries: physical and digital gateways
- Evaluating secondary data: a critical step
What counts as secondary data?
Secondary data is usually understood in contrast to primary data. Primary data comes directly from first-hand sources: your own surveys, interviews, focus groups, or observations. Secondary data, by contrast, has already undergone some degree of processing – it can range from raw organizational records and news archives to fully compiled government reports, academic journals, and industry statistics.
There are broadly two categories to keep in mind. Raw secondary data includes organizational databases, newspapers, and website records – material with little processing. Compiled secondary data covers government publications, books, journals, and research reports that have been selected, summarized, and structured. A third category sits in between: survey-based data such as census records, continuous government surveys on labor markets or household spending, and large-scale attitudinal surveys conducted by independent research organizations.
Understanding these distinctions matters because it helps you select the right kind of data for your specific research question, rather than using whatever shows up first in a search result.
Why communication researchers must cast a wide net
Communication research is inherently interdisciplinary. A study on health messaging draws from public health and psychology. Research on political advertising touches on political science and behavioral economics. An analysis of media literacy in schools overlaps with education policy. This means your secondary data search should never stay within a single discipline.
As the University of Southern California’s research guide on literature reviews rightly notes, thinking about research problems from multiple disciplinary angles is a core strategy for finding new solutions – particularly in the social sciences. Consulting databases across different fields often uncovers perspectives that a single-discipline search would miss entirely. For a communication researcher, this could mean exploring psychology databases for audience behavior studies, education databases for media literacy research, or public health repositories for health communication campaigns.
The internet as a secondary data source
The open web is often the first stop for any researcher, and it can be genuinely useful – but it requires careful navigation. The sheer breadth of information online can be unmanageable, making it easy to waste time sifting through unreliable content. A few practices help manage this.
Using search operators effectively
Rather than typing broad questions into a search engine, use specific operators to filter results. Typing site:.edu restricts results to educational institutions, while filetype:pdf often surfaces research reports and white papers directly. These small adjustments can save significant time and push higher-quality results to the top.
The Internet Archive and digital history
For journalism researchers specifically, historical web content is a legitimate and valuable secondary source. The Internet Archive’s Wayback Machine allows researchers to access past versions of websites, making it possible to track how news coverage or media framing has changed over time. This is particularly useful for longitudinal content analysis – a method that studies patterns in recorded communication using existing texts, widely used in both communication research and historical analysis.
Academic databases: the deep ocean of secondary data
If general internet searches are the shallow end of the research pool, academic databases are where serious research happens. These platforms index peer-reviewed journals, books, conference papers, dissertations, and reports – material that has been evaluated by subject experts before publication. For communication research, several databases stand out.
Google Scholar
Google Scholar is often the best starting point for any literature search. It indexes an enormous range of academic material across all disciplines and is freely accessible without an institutional login. Its citation-tracking feature is particularly useful: you can identify a foundational study in your area and then trace all subsequent work that has cited it, helping you map the conversation within a field over time. Its main limitation is that quality control is not as rigorous as curated databases – not everything indexed is peer-reviewed – so results need to be verified carefully.
JSTOR
JSTOR (Journal Storage) is a digital library of academic journals, books, and primary sources founded in 1994, originally conceived to help research libraries manage the growing volume of academic journals. It is particularly strong in the humanities and social sciences, making it directly relevant to communication research. Unlike Google Scholar’s broad web crawl, JSTOR is a carefully curated archive, focused on reliability and long-term preservation of scholarly content. Access is typically through institutional subscriptions, though registered users can read a limited number of articles free each month. For deep dives into specific historical or theoretical arguments in communication studies, JSTOR is indispensable.
ERIC
ERIC (Education Resources Information Center) is sponsored by the U.S. Department of Education and is the premier database for education-related research. Its relevance to mass communication may not be immediately obvious, but consider this: a large part of communication research intersects with media literacy, the pedagogy of journalism, and how different audiences consume and process information. ERIC provides coverage of journal articles, conferences, government documents, theses, dissertations, reports, audiovisual media, and monographs – making it a rich source whenever your research touches on education, learning, or information processing. If you are studying how school children engage with news media or how journalism is taught across different countries, ERIC is the database to start with.
Communication & Mass Media Complete (CMMC)
For research squarely within the discipline, Communication & Mass Media Complete, available through EBSCO, provides indexing for hundreds of periodicals – including full text for over 200 titles – covering communication and mass media specifically. It features a specialized subject thesaurus with thousands of preferred terms, which helps you find the most relevant academic vocabulary for your search queries. This is particularly useful when you are new to a research area and unsure of the exact terminology used in the literature.
Scopus and Web of Science
For broader, citation-indexed searches across disciplines, Scopus and Web of Science are the two dominant commercial databases. Scopus covers over 90 million core records and additionally provides academic journal rankings and author profiles, while Web of Science covers approximately 100 million items across all major disciplines. Both are typically available through university subscriptions. When you need to verify the impact and credibility of a source – or find who the leading researchers in a particular area are – these two platforms are the most reliable tools available.
Official publications and government data
When your research requires hard numbers – population demographics, literacy rates, internet penetration, media consumption statistics – official sources are non-negotiable. Governments and international bodies are among the largest producers of secondary data in the world, and most of it is freely available to the public.
In India, the National Sample Survey Office (NSSO) under the Ministry of Statistics and Programme Implementation conducts large-scale household surveys that provide data critical for media planners, advertisers, and communication researchers studying rural-urban information access. At the international level, the United Nations provides data on literacy and human rights, the World Health Organization maintains datasets essential for health communication research, and UNESCO publishes extensive reports on media development, journalist safety, and freedom of expression worldwide. The OECD.Stat platform provides social and economic indicators for member and selected non-member economies – useful when comparing media environments across countries.
These sources carry institutional authority and are regularly updated, making them far more reliable for statistical claims than most content found through general web searches.
Libraries: physical and digital gateways
University and public libraries remain among the most underused secondary data resources available to students. Many of the databases discussed above – JSTOR, LexisNexis, Scopus – carry substantial subscription costs that institutions pay on behalf of their students. Logging in through your university library portal typically grants access to millions of documents at no personal cost.
Beyond digital subscriptions, physical library collections sometimes hold materials that have never been digitized: rare manuscripts, archived local newspapers, personal papers of journalists or media figures, and special collections related to regional media history. If you are researching the history of a specific publication or a local broadcasting institution, these physical archives may be the only place to find the evidence you need.
Libraries also employ subject librarians – specialists who can guide you toward the most appropriate databases for your specific research question and help you refine your search strategy. This is a resource that researchers at every level, from undergraduates to doctoral students, consistently underuse.
Evaluating secondary data: a critical step
Finding secondary data is only half the task. Evaluating it rigorously is just as important. Ease of access does not signify credibility – secondary data can be outdated, methodologically flawed, or simply not applicable to your research question. Before using any source, consider its publication date, the credibility of the publisher or sponsoring organization, the methodology behind any statistics or surveys cited, and whether the data was originally collected for a purpose compatible with your research aims.
A useful framework for this evaluation is the CRAAP test – assessing a source’s Currency, Relevance, Authority, Accuracy, and Purpose. Applying this consistently ensures that your research foundation is solid, not just extensive.
It is also worth remembering, as noted in research published in NIH’s PMC, that secondary data was originally collected to answer a different research question. This means quantity of data does not equal appropriateness – a dataset may be vast and credible but still not fit your specific inquiry without careful adaptation.
What do you think? With so many secondary data sources available – from Google Scholar and JSTOR to government databases and library archives – how do you decide which sources to prioritize for a specific research question? And given that communication research is inherently interdisciplinary, do you think researchers in this field are doing enough to draw on data from outside their immediate discipline?
References
- https://methods.sagepub.com/ency/edvol/the-sage-encyclopedia-of-communication-research-methods/chpt/secondary-data
- https://libguides.usc.edu/c.php?g=234974&p=1559473
- https://web.archive.org/
- https://www.scribbr.com/methodology/secondary-research/
- https://scholar.google.com/
- https://www.jstor.org/
- https://en.wikipedia.org/wiki/JSTOR
- https://deepterm.tech/blog/google-scholar-vs-jstor-mastering-academic-research-for-college-success
- https://eric.ed.gov/
- https://libguides.usc.edu/c.php?g=234974&p=1562019
- https://guides.lib.utexas.edu/JOU/articles
- https://www.scopus.com/
- https://www.webofknowledge.com/
- https://paperpile.com/g/academic-research-databases/
- https://mospi.gov.in/
- https://data.un.org/
- https://www.who.int/data
- https://en.unesco.org/themes/media-and-information-literacy
- https://stats.oecd.org/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC7520737/
Leave a Reply