The Research Workflow Problem in PhD Education
PhD programs teach content — the field's concepts, theories, findings, and methodologies — but rarely teach research workflows systematically. Most candidates develop their research process through informal observation, peer advice, and trial and error. The resulting workflows are often either too narrow (searching only in a familiar database, missing important literature in adjacent fields) or too broad (searching everything, reading indiscriminately, never synthesizing to an actionable contribution).
The scholarly ideal — reading everything relevant, synthesizing it comprehensively, identifying the precise gap your dissertation fills, and then designing and executing the empirical or theoretical work to fill it — is not achievable in the 4-7 years available for most PhDs. The practical wisdom is: systematic enough to find what matters and not miss what's important; targeted enough to actually finish.
Research workflows for PhD candidates are the systems and processes that make this possible: systematic literature search strategies that reduce the probability of missing important work; structured approaches to different types of academic research (systematic review, qualitative fieldwork, quantitative empirical study, theoretical contribution); and the documentation practices that make research traceable, reproducible, and defensible at the dissertation defense.
The PhD Research Workflow, Stage by Stage
Stage 1: Literature Search — The Systematic Approach
A literature search for a PhD dissertation is not a single database query. It is a multi-database, multi-strategy search designed to be comprehensive enough to support the claim that you know the relevant literature, have identified its gaps, and are making a genuine contribution.
Multi-database coverage by discipline:
| Discipline | Primary Databases | Secondary/Cross-Disciplinary |
|---|
| Social sciences | JSTOR, ProQuest Social Sciences, Sociological Abstracts | Web of Science, SSRN |
| Psychology | PsycINFO, PsycARTICLES | PubMed (overlap with clinical), CINAHL |
| Education | ERIC, ProQuest Education | SSRN, Google Scholar |
| Business / Management | Business Source Complete (EBSCO), ABI/INFORM | SSRN, Google Scholar |
| Political science | JSTOR, Political Science Complete, ProQuest Political Science | SSRN |
| Economics | EconLit, SSRN (working papers) | Google Scholar, JSTOR |
| Biology and life sciences | PubMed/MEDLINE, Web of Science | Scopus |
| Physical sciences / Engineering | Web of Science, Scopus, IEEE Xplore | arXiv (preprints) |
| Computer science | ACM Digital Library, IEEE Xplore | arXiv, Google Scholar |
| Humanities | JSTOR, Project MUSE, MLA International Bibliography | Google Scholar |
Running the same search strategy across multiple databases is essential because coverage differs substantially by discipline, journal, and publication type. A comprehensive literature search that uses only one database will miss a significant fraction of the relevant work.
Search strategy construction:
Before running searches, construct a formal search strategy: a documented combination of search terms, Boolean operators, and database-specific filters that can be reported in the methodology section of a systematic review or dissertation. The format typically includes:
- Population terms: terms for the population of interest (e.g., "urban renters" OR "residential tenants" OR "housing market participants")
- Concept terms: terms for the main theoretical concept (e.g., "gentrification" OR "residential displacement" OR "neighborhood change")
- Method terms (for systematic reviews): terms to limit to specific study designs (e.g., "longitudinal" OR "panel study" OR "quasi-experimental")
Combined: (urban renters OR residential tenants OR housing market participants) AND (gentrification OR residential displacement OR neighborhood change)
Document this search string and run it in each database. Record the number of results from each database. This documentation is part of the PRISMA flow diagram (or equivalent) that systematic reviews require, and it demonstrates rigor in non-systematic reviews as well.
Citation chaining:
Two systematic literature search techniques that no database query captures:
Backward citation chaining: Find the references in a key paper and check them all. If Cohen & Levinthal (1990) is fundamental to your field, read the papers it cites — these are the foundational literature that Cohen & Levinthal built on, and you need to read them to understand the theoretical lineage.
Forward citation chaining: Using Google Scholar's "Cited by" function or Web of Science's "Cited References" search, find all papers that have subsequently cited a key paper. If you've identified Cohen & Levinthal (1990) as foundational, the papers that cite it are the subsequent literature that built on, extended, critiqued, or applied the foundational work. This is often how you find the most current and active debates in a field.
Both directions of citation chaining should be applied to every paper you identify as genuinely foundational to your dissertation.
Grey literature:
Published academic journal articles are not the only relevant source for most dissertation research. Depending on your field and topic:
- Working papers (SSRN, NBER, IZA, CESifo) — often more current than published papers; may represent where the field is heading
- Government reports and policy papers (CBO, GAO, think tanks, agency research divisions)
- Conference papers (SSRN, conference proceedings, institutional repositories)
- Institutional reports (NGO reports, industry association research, foundation publications)
- Dissertations (ProQuest Dissertations & Theses) — especially relevant for identifying who else has worked on your topic and how they approached it
Grey literature requires quality assessment because it lacks peer review. But dismissing it entirely risks missing significant empirical work, policy-relevant analyses, and working papers that will eventually be the field's next landmark publications.
Stage 2: Literature Synthesis — From Reading to Contribution Claim
The point of reading the literature is not to have read it — it is to develop a precise, defensible contribution claim: what does your dissertation add to what the literature has already established?
The path from "I've read the literature" to "here is my specific contribution claim" involves several synthesis steps that most PhD programs teach too late:
Conceptual mapping:
As you accumulate literature, organize it visually: which theoretical positions are there? Which empirical findings support which theoretical claims? Where do findings conflict, and why (different populations? different operationalizations? different time periods?)? What domains or populations are conspicuously absent from the existing literature?
This mapping can be literal (a diagram on paper, or a concept map tool like CmapTools), or structural (Zotero collections organized by theoretical cluster), or both. The goal is a comprehensive view of the field's landscape — not the sequential reading list, but the simultaneous intellectual map.
Gap analysis:
The contribution claim emerges from the gap analysis. Gaps in the literature fall into several types:
- Population gap: The phenomenon has been studied in population X but not population Y. (e.g., absorptive capacity has been extensively studied in large firms but rarely in small and medium enterprises.)
- Context gap: The phenomenon has been studied in context X but not context Y. (e.g., organizational resilience literature is primarily from the US/UK; how does it apply in the organizational context of East Asian firms?)
- Mechanism gap: We know A causes B, but the mechanism is unclear. (e.g., we know social capital predicts income, but the specific pathways — through information, through job referrals, through access to finance — are not well-identified.)
- Methodological gap: Existing evidence comes from cross-sectional studies; no longitudinal evidence exists. (e.g., the causal claim requires panel data, not cross-sectional surveys.)
- Theoretical gap: Two theoretical traditions have developed in parallel without integration; they would illuminate each other. (e.g., institutional theory and resource-based view both predict firm behavior but rarely speak to each other directly.)
Your contribution claim is: "This dissertation addresses the [type] gap in the literature by [approach], using [method] to establish [finding or argument]."
Stage 3: Research Design — The Empirical or Theoretical Architecture
Once the contribution claim is established, the research design follows from it: what evidence or argument would establish the contribution you're claiming?
For qualitative empirical research:
Qualitative methods — ethnography, in-depth interviews, case studies, grounded theory, discourse analysis — are appropriate when the research question concerns the meaning, process, or mechanism of a phenomenon rather than its frequency or magnitude. The research design decision is: which qualitative method is appropriate given your research question and access to the phenomenon?
Key design decisions in qualitative research:
- Site or participant selection: How will you access the phenomenon? Who or what are the units of analysis?
- Sampling strategy: Theoretical sampling (selecting based on theoretical relevance, continuing until theoretical saturation), purposive sampling, or snowball sampling for hard-to-access populations
- Data collection: Interviews (structured, semi-structured, or unstructured?), observation, document analysis, or some combination
- IRB protocol: All research involving human subjects requires IRB approval (see Compliance section below). The IRB process must be completed before data collection begins.
For quantitative empirical research:
Quantitative methods — surveys, experiments, regression analysis, natural experiments, panel data analysis — are appropriate when the research question concerns the distribution, magnitude, or causal effect of a phenomenon.
Key design decisions in quantitative research:
- Data source: Primary data collection (survey, experiment) or secondary data analysis (existing datasets, administrative records)?
- Causal identification strategy: If the goal is causal inference, what is the identification strategy? Random assignment (experiment), instrumental variables, difference-in-differences, regression discontinuity, synthetic control? Observational data without a causal identification strategy can establish correlations but not causation — this is a limitation that must be acknowledged clearly.
- Measurement: How will you operationalize the concepts you're studying? What reliability and validity evidence supports the measurement approach?
- Sample size and power: A priori power analysis determines the minimum sample needed to detect an effect of a given size at a given significance level. Running a study without adequate power risks both Type II errors (missing real effects) and (when sample sizes are post-hoc selected) inflated effect estimates.
For theoretical contributions:
Not all dissertations are primarily empirical. Some contribute theoretical frameworks, conceptual integrations, or formal models. The research design for theoretical contributions involves: precise statement of what the new framework or model is, how it differs from existing theories, what phenomena it explains that existing theories cannot, and what empirical predictions it makes.
Stage 4: Documentation, Reproducibility, and Pre-Registration
Pre-registration (for empirical research):
Pre-registration — documenting your hypotheses, research design, and analysis plan before data collection — is an increasingly expected practice in psychology, economics, and other empirical social sciences. It distinguishes confirmatory research (testing pre-specified hypotheses) from exploratory research (developing hypotheses from data), prevents p-hacking and outcome switching, and substantially increases the credibility of findings. The Open Science Framework (OSF, osf.io) provides free pre-registration infrastructure; the AEA RCT Registry is the standard for randomized experiments in economics.
Data management plan:
Funding agencies (NSF, NIH, SSHRC) increasingly require data management plans specifying how data will be collected, stored, protected, and eventually shared (or why sharing is not feasible). Many universities require a data management plan as part of dissertation approval. The key components: data collection procedures, storage and backup (encryption if sensitive), data sharing plan (will data be deposited in a public repository?), and plan for human subjects data (retention requirements, de-identification before sharing, IRB-specified protocols).
A Recommended Tool Stack for PhD Research Workflows
| Stage | Tool | Notes |
|---|
| Database search | Web of Science, Scopus, JSTOR, PsycINFO | Discipline-specific; run searches on multiple |
| Citation chaining | Google Scholar ("Cited by") | Free; enables forward citation chaining |
| Working papers | SSRN, NBER, arXiv | Free preprints; often 1-2 years ahead of publication |
| Citation management | Zotero | Free; import directly from most databases |
| Grey literature access | ProQuest Dissertations, govt websites | Dissertations; government reports |
| Pre-registration | OSF.io, AEA RCT Registry | Free; increasingly expected |
| Data storage | OSF.io, institutional repository | Version control; sharing-ready |
| Web resource capture | WebSnips | Government reports, grey literature, working papers |
WebSnips for PhD research workflows: Government statistical data, policy documents, and grey literature on institutional websites are essential sources in many research areas — and are among the highest-risk sources for link rot. A study published in PLOS ONE in 2013 found that a significant proportion of URLs in academic publications become inaccessible within a few years; for government and institutional URLs, the rate is often higher. WebSnips captures these sources with date and source URL at the time of access, creating a retrievable, dated archive. For a PhD candidate working on housing policy research, capturing the relevant HUD reports, Census Bureau housing datasets, and state housing authority publications with WebSnips provides both link rot protection and an organized, dated evidence trail — the capture date is academically significant for data that is periodically updated (Census tables, BLS statistics, agency guidelines). Organized by research project (Dissertation Chapter 2 Primary Sources, Policy Document Archive), WebSnips clips build the grey literature layer of the research workflow.
A Worked Example: Systematic Literature Search for a Sociology Dissertation
A sociology PhD candidate, Alex Rivera, is writing a dissertation on how online platforms are reshaping gig work. She has developed her research question: "How do algorithmic management systems in gig economy platforms shape the autonomy, identity, and resistance strategies of platform workers?"
Step 1 — Search strategy construction:
Alex documents her search strings before running any searches:
Population terms: gig workers OR platform workers OR freelance workers OR crowdworkers OR on-demand workers
Concept terms: algorithmic management OR algorithmic control OR platform governance OR digital labor
Context terms: gig economy OR platform economy OR sharing economy OR app-based work
Combined string: (gig workers OR platform workers OR freelance workers OR crowdworkers) AND (algorithmic management OR algorithmic control OR platform governance) AND (autonomy OR identity OR resistance OR worker experience)
Step 2 — Multi-database search execution:
Alex runs this search in: JSTOR (1,847 results, filtered to 2015-2026, yields 412), Sociological Abstracts (183), Business Source Complete (291), and Google Scholar (broad sweep, cross-checked for missing key papers). She uses inclusion/exclusion criteria to screen titles and abstracts: English-language; peer-reviewed; focus on platform workers specifically (not gig work in non-platform contexts); at least some empirical content.
After screening, she has 64 papers for full-text review.
Step 3 — Citation chaining:
From the four most-cited papers in her initial search, she runs forward and backward citation chaining. This adds 18 papers she hadn't found through keyword search, including three foundational pieces on algorithmic management in non-gig contexts (Amazon warehouse, Uber driver, and content moderation) that inform the theoretical framework.
Step 4 — Grey literature:
Alex searches ProQuest Dissertations for recent doctoral theses on algorithmic management (finds two highly relevant completed dissertations; reads their bibliography sections). She also captures three relevant reports from Worker Information Exchange and Data & Society Research Institute via WebSnips — both organizations publish empirical research on gig work that doesn't appear in peer-reviewed databases.
Step 5 — Synthesis and gap identification:
After reading, Alex creates a conceptual map. She identifies that: (1) most existing research focuses on Uber and Lyft specifically; (2) resistance strategies are understudied relative to control mechanisms; (3) no existing study has used ethnographic methods to observe platform worker organizing on the platform (rather than studying organizing that has already occurred). Her contribution: ethnographic study of spontaneous resistance formation among Amazon Mechanical Turk workers, addressing the methodological gap and the resistance gap simultaneously.
Compliance and Research Ethics
IRB/Ethics board approval:
Any research involving human subjects — interviews, surveys, ethnography, observation — requires ethics review and approval before data collection begins. In the US, this is the Institutional Review Board (IRB); in Canada, the Research Ethics Board (REB); in the UK, the ethics review process varies by institution. IRB review determines whether human subjects protections (informed consent, privacy protection, data security, minimization of risk) are adequately addressed.
Informed consent:
Research participants must give informed consent before participating in any research that involves them. Consent must be voluntary, informed (participants must understand what they're consenting to), and documented. Research on vulnerable populations (minors, incarcerated individuals, cognitively impaired individuals) requires additional protections.
Data security:
Human subjects data must be stored securely. De-identification standards depend on the nature of the data and the IRB protocol. Identifiable human subjects data should never be stored in unsecured personal cloud storage without IRB approval of the storage method.
Plagiarism and self-plagiarism:
Academic plagiarism — using another's words or ideas without attribution — is a career-ending violation. Self-plagiarism — recycling your own prior work without disclosure — is also a violation and is taken seriously by journals and dissertation committees. Cite all sources, including your own prior publications if you incorporate them.
Common PhD Research Workflow Mistakes
Mistake 1: Single-database literature searches.
No single database covers even a majority of the relevant literature in most fields. A Scopus-only search misses JSTOR; a JSTOR-only search misses working papers on SSRN; a search without grey literature misses policy reports and dissertations. Run searches across multiple databases.
Mistake 2: Not documenting the search strategy.
An undocumented literature search cannot be replicated, updated, or defended. Document every search: database, date, search string, number of results, inclusion/exclusion criteria applied. This documentation is required for systematic reviews and valuable for all dissertation research.
Mistake 3: Skipping forward citation chaining.
Finding the papers that cite your key papers often surfaces the most current, active research in your area — precisely the literature that will be most relevant to the scholarly conversation your dissertation is entering. Google Scholar's "Cited by" function takes 5 minutes per paper and can surface literature that keyword searches miss.
Mistake 4: Starting data collection before IRB approval.
Collecting data before IRB approval is a serious research ethics violation. If the violation is discovered, the data typically cannot be used in the dissertation. Begin the IRB process as soon as the research design is finalized — approval timelines range from a few days (exempt research) to several months (higher-risk protocols).
Key Takeaways
- Research workflows for PhD candidates require systematic multi-database literature searches, documented search strategies, citation chaining in both directions, grey literature inclusion, and formal gap analysis — not a single database query.
- Document every search: database, date, search string, results count, inclusion criteria — this documentation is required for systematic reviews and valuable for all dissertation defense situations.
- Citation chaining in both directions: backward chaining finds foundational literature; forward chaining (Google Scholar "Cited by") finds current debates. Both are essential for comprehensive literature coverage.
- Gap analysis produces the contribution claim: the dissertation contribution is identified by mapping the conceptual landscape of the field and identifying the specific population, context, mechanism, methodological, or theoretical gap your work addresses.
- Pre-register before data collection: for empirical research, pre-registration on OSF or an equivalent registry protects against bias accusations and substantially increases the credibility of findings.
- IRB approval must precede data collection: this is not optional; collecting data before IRB approval is a research ethics violation with serious consequences.
Conclusion
Research workflows for PhD candidates are the systematic practices that turn the vast landscape of academic knowledge into a navigable path from literature to contribution. The PhD candidates who develop these workflows early — multi-database searches with documented strategies, citation chaining, grey literature coverage, formal gap analysis, and pre-registration before empirical work — produce more defensible dissertations, avoid the most common research integrity problems, and develop the scholarly habits that support a career of contribution rather than just one dissertation. The research process is not a hurdle before the writing; it is the intellectual substance of the dissertation itself.
Try WebSnips free — clip government data reports, policy documents, grey literature, working papers, and institutional publications with date and source URL, building the dated, archived research source library that complements systematic database searches and protects PhD research against link rot throughout the dissertation process.