Click here for the latest version

JEL Classification: A10, B41, C18, C80, D83

Introduction

Economic research has undergone a significant transformation over the past few decades, marked by an increased emphasis on establishing causal relationships through empirical methods. This "credibility revolution" has propelled the discipline toward more rigorous identification strategies, aiming to provide robust evidence for policy-making and theoretical advancement ([8]). Pioneering work by researchers such as Orley Ashenfelter, Joshua Angrist, David Card, Guido Imbens and Alan Krueger introduced methodologies that enhanced causal identification, including natural experiments, regression discontinuity designs (RDDs), and instrumental variables (IVs) ([9, 40]). These approaches address endogeneity concerns and provide more credible estimates of causal effects.

Leading journals now prioritize studies employing these methods over traditional correlational approaches ([20, 37]). Hamermesh (2013), for instance, documents a decline in purely theoretical articles and an increase in empirical studies utilizing self-generated data and experimental methods in top economics journals.3 The credibility revolution has raised the bar for empirical work, emphasizing transparent reporting, careful consideration of identification assumptions, and rigorous sensitivity analyses ([7, 40]). However, this focus on specific methodologies has sparked debates about the direction and priorities of economic research.4 Studies like Backhouse and Cherrier (2017); Currie et al. (2020); Goldsmith-Pinkham (2024) have also shown a significant rise in empirical methods across economics, supported by advancements in data and technology. Our study further contributes by breaking down these trends into specific causal inference methods and examining their usage across subfields.

Despite extensive discussions on methodological advancements, there is a lack of comprehensive analysis that quantifies how the structure and complexity of economic research have evolved over time, particularly regarding the use of causal claims and narrative complexity across subfields. Our study addresses this gap by analyzing over 44,000 NBER and CEPR working papers using a custom large language model to extract structured information on the knowledge graphs of papers, including the methods used to evidence each claim—causal or otherwise—and the data employed in the analyses.

We introduce a novel approach by constructing a knowledge graph for each paper in our dataset. In these graphs, nodes represent economic concepts classified using JEL codes, and edges represent relationships from a source node to a sink node. This means that if a paper discusses how one economic concept relates to another, we capture this as a directional link between those concepts. Whether or not a claim is considered causal depends on the method used to substantiate it. Specifically, we identify an edge as a causal edge if the claim is evidenced using causal inference methods such as Difference-in-Differences (DiD), Instrumental Variables (IV), Randomized Controlled Trials (RCTs), Regression Discontinuity Designs (RDDs), Event Studies, or Synthetic Control

To systematically evaluate these knowledge graphs, we develop three broad categories of measures. First, we track the narrative complexity of a paper, including the breadth and depth of claims. Second, we examine novelty and contribution, capturing whether a paper’s relationships are genuinely new or whether it “fills gaps” previously underexplored in the literature, distinguishing between causal and non-causal contributions. Third, we consider conceptual importance and diversity, focusing on how centrally a paper’s concepts sit within the overall network of economic ideas and whether the paper balances multiple causes (sources) with multiple outcomes (sinks). Additionally, by distinguishing these measures based on claims in the non-causal subgraph from those in the causal subgraph, we parse out the difference between general narrative features and those supported by rigorous identification.

Our analysis reveals several notable patterns. First, the use of causal inference methods has significantly increased over time, with the average proportion of causal claims rising from about 4% in 1990 to nearly 28% in 2020, reflecting the impact of the credibility revolution in economics.

Second, the determinants of publication and citation success differ markedly. While employing rigorous causal inference methods, introducing novel causal relationships, and engaging with frontier topics robustly increase the likelihood of publication in top 5 journals, these same features do not necessarily translate into higher citation counts post-publication. Instead, papers focusing on central, widely recognized concepts accumulate more citations, highlighting a divergence between editorial preferences for causal innovation and the broader academic audience’s sustained attention.

Third, the complexity of a paper’s causal narrative—such as the depth of its longest causal chain or the number of distinct pathways linking causes and effects—plays a pivotal role in both publication and citation outcomes. Causal narrative complexity is particularly associated with top 5 publications and higher citation counts. This suggests that “deeper” causal arguments can succeed at both gatekeeping and long-run academic influence. In contrast, non-causal narrative complexity is generally uncorrelated or negatively associated with academic success.

These findings highlight the trade-offs between methodological rigor, novelty, topic centrality, and broader research recognition. While top-tier journals prioritize innovative and methodologically sophisticated research, citation impact often depends more on engaging with central and well-established concepts. This tension underscores the evolving dynamics of academic evaluation in economics, where causal rigor and novelty increasingly shape short-term outcomes, but long-term impact may still rely on established conceptual foundations.

Literature Critics argue that the emphasis on specific empirical methods and complex narratives may lead to overconfidence in results and potential overinterpretation, especially if the underlying assumptions are not fully met ([28]; [44]).5 Additionally, the focus on identification sometimes comes at the expense of economic theory, resulting in studies that establish causal effects without adequately explaining the underlying mechanisms ([62]; [38]). As [44] notes, the detachment from theoretical frameworks can limit the explanatory power of empirical findings. This underscores that while methodological rigor is essential for credible causal inference, it should not preclude the consideration of valuable evidence from diverse sources. Misapplication or overinterpretation of methods can lead to questionable conclusions. For instance, the use of instrumental variables relies on strong assumptions that the instrument affects the outcome only through the endogenous explanatory variable and is uncorrelated with the error term ([9]). Violations of these assumptions, such as weak instruments or invalid exclusion restrictions, can produce biased estimates ([58]). For example, [48] highlights the challenges of using weather variables as instruments, identifying numerous potential exclusion-restriction violations.

RCTs have gained prominence as a gold standard for causal inference due to their internal validity. However, scholars like [29] and [21] argue that RCTs may suffer from limited external validity and may not capture complex economic phenomena, and can further be subject to inducing demand effects ([27]). Generalizing findings from specific experimental settings without considering contextual differences can lead to misleading conclusions ([55]).

Moreover, the complexity of research outputs has increased as papers have become longer and include more coauthors ([20]).6 This expansion reflects the need to address methodological rigor and to include detailed explanations of causal mechanisms, robustness checks, and theoretical integration. However, this rise in narrative complexity may also indicate that the presentation and promotion of research findings are becoming more important factors in dissemination.7 Increased complexity can make it challenging for readers and reviewers to critically assess the validity of the claims ([41]; [34]), and may contribute to overemphasizing research findings. The "garden of forking paths" metaphor illustrates how analytical flexibility can lead to false-positive findings even without intentional misconduct ([34]).8 The American Statistical Association has highlighted the misinterpretation and misuse of p-values, advocating for a more nuanced understanding of statistical significance ([66]). [61] propose the p-curve method as a tool to detect and correct for publication bias using only significant results, highlighting the pervasiveness of selective reporting.

The relationship between theory and empirics is a central concern in economics. Critics argue that the focus on empirical identification has led to a neglect of theoretical development ([44, 38]). Without a solid theoretical foundation, empirical findings may lack coherence. [62] emphasizes that economics is not purely an experimental science and that theoretical models are essential for interpreting empirical results. Furthermore, [5] highlight that economists see value in research that is multidisciplinary and addresses diverse topics, suggesting a need to balance empirical rigor with theoretical and interdisciplinary approaches.

Finally, emerging methodologies such as machine learning, Bayesian inference, and the increased reliance on big data pose further challenges in training, replication, and interpretation.9 Amid these technical shifts, the ethical and communicative dimensions of research are also in flux, with social media and open-data debates reshaping public trust ([57, 3, 33]).

Ethical considerations in empirical research extend beyond methodological rigor to include transparency about limitations, uncertainties, and the broader context of findings ([57]). Misleading claims can distort policy-making, erode public trust, and lead to poor allocation of resources. Ensuring integrity in research is a collective responsibility involving researchers, journals, institutions, and funding bodies.10 Moreover, the ways in which academics engage with the public—particularly through social media—can shape the perceived credibility of their work.11

These concerns gain even more urgency in light of recent global initiatives emphasizing the importance of evidence-based policy-making. International leaders and organizations have expressed alarm over the slow progress in achieving Sustainable Development Goals (SDGs), attributing part of the challenge to a lack of robust evidence to inform policy decisions. Despite substantial investments in public services, a "hidden" repository of underutilized studies exists that could inform better policy choices ([51]). Recognizing these challenges, research councils and governments are investing in innovative solutions to enhance the accessibility and synthesis of existing research. For instance, in September 2024, the UK Economic and Social Research Council (ESRC) announced a significant investment in artificial intelligence to facilitate evidence synthesis for public policy, aiming to build a global infrastructure that provides useful evidence for policymakers.12 Moreover, organizations like the Behavioural Insights Team have proposed blueprints for better international collaboration on evidence, emphasizing the importance of evidence synthesis and accessibility ([15]).13

Understanding these trends requires examining how economic research evolves over time. [6] analyze a large dataset of economics journal articles from 1980 to 2015, documenting shifts in research fields and styles, with more empirical papers appearing in influential journals and receiving more citations. This evolution underscores the growing methodological sophistication in economics but also raises concerns that genuinely new ideas may be harder to find ([16], [53]). Our analysis shows that journals sometimes reward gap filling papers—whether bridging underexplored topics pairs or linking distinct subfields—yet these connections do not always translate into greater long-run citation impact. That said, top-tier outlets particularly value interdisciplinary bridging when it involves causal claims, suggesting an appetite for innovative cross-field insights that are firmly grounded in causal inference methods.14

Our study contributes to this discussion by providing empirical evidence on the evolution of empirical methods and their differential adoption across subfields. We observe that methods such as DiD, IV, RCTs, and RDDs have seen substantial growth, reflecting the discipline’s shift toward more rigorous identification strategies. Fields such as Urban, Health, Development, and Behavioral exhibit the most significant increases in the use of causal inference methods. Conversely, fields like Macroeconomics show more modest growth. This variation highlights how research questions and methodological traditions influence the adoption of empirical methods across different areas of economics.

By constructing and analyzing knowledge graphs of economic research, our study offers a new perspective on how the complexity and structure of narratives have changed and how these features shape both journal acceptance and scholarly recognition. The remainder of this paper is organized as follows: Section 2 details the data and information retrieval methods used to construct these graphs. Section 3 documents the rise of causal inference and introduces our three categories of measures—narrative complexity, novelty and contribution, and conceptual importance and diversity—examining their evolution and cross-field variation. Section 6 then links these measures to publication outcomes and citation impacts. Section 7 concludes with broader implications for research practices, methodological priorities, and the communication of economic knowledge in an era defined by the credibility revolution.

Data and Methods

In this section, we present the data sources, extraction processes, and methods employed to examine causal claims within the economics literature. Further methodological details and technical specifications are provided in Appendix A.15

Working Paper Corpus

Our analysis is based on a comprehensive corpus of working papers from two primary sources: the National Bureau of Economic Research (NBER) and the Centre for Economic Policy Research (CEPR). The NBER dataset comprises 28,186 working papers, while the CEPR dataset includes 16,666 papers, resulting in a total sample of 44,852 papers. These papers span several decades and encompass various subfields of economics, providing a broad view of the research landscape.

To refine the sample and focus on relevant content, we applied specific filtering criteria. We included only papers containing more than 1,000 characters to exclude incomplete documents. Additionally, we limited the analysis to the first 30 pages of each paper, ensuring that we captured the sections most likely to contain causal claims, such as introductions, literature reviews, and empirical analyses.

The corpus covers a wide range of economics subfields, including Labour Economics, Public Economics, Macroeconomics, Development Economics, and Finance. The papers employ diverse empirical strategies, such as Randomized Controlled Trials (RCTs), Instrumental Variables (IV), Difference-in-Differences (DiD), and Regression Discontinuity Designs (RDD), allowing us to examine methodological trends across the discipline.

Pre-processing The preprocessing of the text data followed a structured pipeline aimed at cleaning and normalizing the text for analysis. The preprocessing steps included removing excessive whitespace, converting all characters to lowercase, and filtering out non-alphanumeric characters, keeping only spaces for readability. Additionally, we stripped leading and trailing whitespace to ensure uniformity in the text. These steps were helpful for the large language model to efficiently and accurately process the text, especially considering the large volume of data involved.

LLM based retrieval

We employed a multi-stage process using a large language model (LLM) to extract and analyze information from the working papers in our corpus.16 We interacted with the LLM using carefully designed prompts that guided the model to extract the required information while adhering to a predefined JSON schema. The overall process is visually summarized in Figure 1, which illustrates the flow from input text to structured data extraction and subsequent analysis. This approach allowed us to efficiently process the text and extract detailed structured data necessary for our analysis while minimizing computational and human resources.17

Retrieval of Concepts Using AI flowchart

Figure 1. Note: This flowchart illustrates our AI-powered approach to retrieving, assessing, and mapping causal claims and contributions from academic papers. The process begins with academic papers, from which the LLM extracts fields such as Author, Publication, Institution, Field, Method, and Data/Code Availability. These aspects feed into two main branches: Identification and Causal Claims. The Identification branch focuses on elements like Identification Strategy and Robustness Checks. The analysis extends to understanding precise measurements and contexts, as well as extrapolated concepts and contexts, leading to insights on contributions claimed and policy recommendations. The Causal Claims branch involves analyzing the causal and non-causal relationships identified in the papers, consisting of arrays of source (or cause) and sink (or effect) variables. The analysis operates across three levels. First, for each source or sink node, we consider the source of sink as claimed by the author and as measured in the paper, including the type the owner of the data used. Second, for each source-sink edge, we examine the method(s) used to evidence a claim, and whether null result was found. Third, at the graph level, we assess graphical measures like the number of steps taken from source to sink, the descriptions of these steps, and the overall complexity of the underlying narrative.

(Figure 1)

Our LLM-based retrieval process consists of the following stages:

Stage 1: Qualitative Summary Extraction In the first stage, we prompted the LLM to analyze the first 30 pages of each paper and extract a curated summary of key elements.18 This included the research questions as presented in the abstract, introduction, and full text; information on causal identification strategies used in the paper; details on data usage, accessibility, and acknowledgements; and metadata such as authors’ names, institutional affiliations, fields of study, and methods used. This initial extraction provided a structured overview of each paper, which was used in subsequent stages to extract more detailed information. To ensure the reliability of our extraction process, we validate our information retrieval methods in Appendix C, demonstrating high accuracy and F1 scores across key empirical methods and fields of study.

Stage 2: Extraction of Causal Claims Using the curated summaries from Stage 1, particularly the sections on causal identification and causal claims, we prompted the LLM to extract detailed all knowledge links presented in each paper between two knowledge entities. The LLM identified source and sink variables as described by the authors, determined the types of relationships (e.g., direct causal effect, indirect causal effect, mediation, confounding, theorised relationship, correlation), and recorded the empirical methods used to establish each link (e.g., RCT, IV, DiD, OLS, simulations). The result was an edge list per paper, where each row represents an edge with a source node (e.g., a cause) and a sink node (e.g., an effect). We also included relevant edge attributes, such as the method used to evidence that edge. While we collected additional attributes like the direction of effect, magnitude, and statistical significance, these features were experimental and are not used in the main analysis due to variation in reporting standards.19

Stage 3: Data Usage and Accessibility Extraction From the data-related summaries in Stage 1, we prompted the LLM to extract structured information regarding data sources and accessibility. Key elements included the ownership of the data (e.g., private company, public sector entity, researchers), data accessibility (e.g., freely accessible, restricted), and details on data granularity, units of analysis, temporal and geographical context. This information is crucial for assessing trends in data usage and the implications for transparency and replicability in economic research.

Validation of Information Retrieval To validate our information retrieval methods, we conducted two exercises (details in Appendix C). First, we matched 307 papers with the annotated dataset from [18], which classifies empirical methods and fields for 1,106 economics papers. Our retrieval achieved high accuracy and F1 scores, especially for RDD methods and fields like Macroeconomics and Urban Economics, demonstrating reliability across key dimensions.

In a second exercise, we compared our causal claims data to the Plausibly Exogenous Galore dataset,20 a source documenting primary causal variables and exogenous variation for 1,435 papers. We matched 485 papers and aggregated our data at the paper level to align with their structure. This comparison yielded moderate similarity for causes and effects, likely due to the Plausibly Exogenous Galore dataset’s focus on the most important causal link, whereas our dataset captures the full knowledge graph of each paper, including all causal links. This broader scope introduces variability in matching specific causes and effects. However, the source of exogenous variation showed high similarity, supporting the consistency of our approach in capturing essential causal elements across datasets.

Matching Variables to Standardized Economic Concepts

To facilitate systematic network analysis and aggregation, we standardized the free-text descriptions of the source and sink variables by mapping them to official Journal of Economic Literature (JEL) codes. We created semantic embeddings for each JEL code’s overall description, which concatenates the JEL description, guidelines, and keywords.21 By generating vector embeddings of the variable descriptions and comparing them to the vector embeddings of JEL code descriptions using cosine similarity, we identified the most relevant codes for each variable.22 This process situates each causal claim within the broader context of economic research and allows us to construct a knowledge graph of economic research, mapping and documenting the frontier in causal evidence over time. This process is visually summarized in Figure 2, which illustrates our AI-driven approach to analyzing and mapping causal linkages between JEL codes. Full details are available in Appendix Section B.

Mapping causal linkages between JEL codes using AI

Figure 2. Note: This diagram illustrates our AI-driven methodology for analyzing and mapping causal and non-causal linkages between economic concepts, represented by JEL (Journal of Economic Literature) codes. Starting with a corpus of working papers, we use a custom prompt and pre-trained language model to extract causal relationships, identifying source (or cause) and sink (or effect) variables within the text. The extracted edge are parsed to generate directed linkages between JEL codes, forming a knowledge graph that aggregates these relationships across the corpus. We employ OpenAI’s vector embeddings to numerically represent descriptions of JEL codes and utilize cosine similarity with sources and sinks, assigning the most similar JEL code to each of the source and sink nodes. This approach enables us to construct a structured representation of evidence in economics over time, facilitating the exploration of interconnected economic concepts and the evolution of empirical research frontiers.

(Figure 2)

Citations and Publication Data

Matching Publication Outcomes to Working Papers To analyze the publication trajectories of the working papers, we matched each paper to its eventual publication outcome using multiple data sources. Our primary goal was to determine whether a working paper was published in a peer-reviewed journal and, if so, identify the journal and publication date. This information is essential for understanding the dissemination and impact of research within the economics discipline.

We utilized four data sources to obtain publication information. First, we used official metadata from the NBER, which provides publication data collected via author submissions and automated scraping from RePEc.23 While comprehensive, the dataset includes duplicates and primarily covers NBER papers. The second source was a large language model (LLM) prompted to retrieve publication outcomes based on its knowledge, yielding results for a small subset of NBER and CEPR papers. Third, we used the OpenAlex repository, matching titles of working papers and prioritizing published versions when multiple matches existed. Finally, we incorporated data from [14] (2020), which provides manually verified publication outcomes for NBER and CEPR papers between 2000 and 2012, matched up to mid-2019.

To ensure consistency, we standardized journal names across these sources using the SCImago Journal Rank (SJR) lists for the fields of "Economics, Econometrics and Finance" and "Business, Management and Accounting." After removing generic journal names, this list included 2,367 unique journals.

Our matching process followed a hierarchical approach, prioritizing verified data. We first checked for a match in the dataset from [14] , followed by a search in OpenAlex for exact title matches. If no match was found, we consulted the NBER metadata, and finally used the LLM retrieval method for remaining papers. This approach ensured a comprehensive and reliable matching of publication outcomes. In total, the dataset from [14] provided publication outcomes for 9,139 papers, OpenAlex matched 10,840 papers, NBER metadata contributed 15,872 matches, and the LLM retrieval identified outcomes for 1,707 papers. By consolidating these sources, we obtained a detailed picture of the publication trajectories of a substantial number of working papers.

Despite this extensive coverage, certain limitations remain. The NBER metadata may be incomplete due to reliance on author submissions, and the [14] dataset only covers papers up to 2012, matched to 2019. Additionally, errors may arise due to title similarities or data entry issues. However, by leveraging multiple sources, we minimized these limitations, resulting in a robust dataset for further analysis.

Citations Data To extend our analysis, we collected citations data for the working papers using three primary sources: RePEc’s CiteEc service (https://citec.repec.org/), [14], and the OpenAlex repository. We prioritized the citations data from RePEc’s CiteEc service, which provides up-to-date (as of November 2024) citation counts for a large number of economics papers. For papers not included in CiteEc, we used the manually verified citations from [14], which provides citation counts for NBER and CEPR papers published between 2000 and mid-2019. For any remaining papers, we obtained citation counts from OpenAlex, matched by exact paper titles. By merging these sources and prioritizing in this order, we assembled citations data for approximately 94.6% of our total sample, and 97.7% of the pre-2020 sample used in our analysis. This extensive coverage enables us to incorporate citations as a measure of research impact.

Graphical Framework

Over the past four decades, economic research has undergone a profound transformation, characterized by an increasing emphasis on establishing credible causal relationships using empirical methods—a shift often referred to as the “credibility revolution.” Scholars such as [7] have highlighted the importance of rigorous econometric techniques designed to enhance causal inference, describing them as Mostly Harmless Econometrics. To systematically capture and analyze this evolution, we introduce a graphical approach by constructing a knowledge graph for each paper in our dataset. This method allows us to quantitatively assess the complexity and structure of claims in economic research over time and observe changes in the adoption of causal inference methods.24

For each paper pp, we define a directed graph Gp=(Vp,Ep)G_p = (V_p, E_p), where VpV_p is the set of nodes representing economic concepts, classified using JEL codes, and EpE_p is the set of directed edges representing claims from a source node to a sink node. We use the terms “source” and “sink” to denote the direction of the claim within the paper, without presupposing causality. Whether a claim is interpreted as causal depends on the attributes of the edge connecting the nodes.

An important attribute of each edge e∈Epe \in E_p is whether the claim was evidenced using a causal inference method. We classify an edge as a causal edge if the associated claim in the paper was evidenced using one of the following methods: Difference-in-Differences (DiD), Instrumental Variables (IV/2SLS), Randomized Controlled Trials (RCTs/Experiments), Regression Discontinuity Design (RDD), Event Study, or Synthetic Control. This classification allows us to distinguish between causal and non-causal claims within the network. With this definition, approximately 19% of all claims in our dataset are classified as causal edges.

The network GpG_p thus includes both causal and non-causal edges. For instance, theoretical relationships between two concepts are represented as edges but are not considered causal unless they are supported by the specified empirical methods. This approach enables us to analyze the overall structure of a paper’s argumentation and the role of causal inference methods within it.

Observing the Credibility Revolution in Economic Research

To analyze the evolution of the use of causal inference methods in economic research, we focus on the proportion of causal edges in papers over time. This measure reflects the extent to which economists have increasingly adopted rigorous causal inference methods in their work, indicative of the credibility revolution.

Figure 3 Figure 3 (a) displays the average proportion of causal edges per paper from 1980 to 2023. The data show a significant increase over time. In 1990, the average proportion of causal edges was approximately 4.2%. By 2000, it had risen modestly to around 8.4%. However, the increase became more pronounced in the subsequent decades: by 2010, the average proportion reached approximately 17.1%, and by 2020, it had climbed to around 27.8%. This upward trend indicates that economic papers have increasingly incorporated causal inference methods to substantiate their claims over the past three decades.

Trends in the Proportion of Causal Edges Over Time and by Field

Figure 3. Note: This figure presents the trends and distribution of the average proportion of causal edges per paper in NBER and CEPR working papers across different dimensions. Panel (a) displays the average proportion of causal edges per paper from 1980 to 2023, showing a significant increase from approximately 4.2% in 1990 to around 27.8% in 2020. The solid blue line represents the average, and the shaded area indicates the 95% confidence interval. Panel (b) shows the average proportion of causal edges by field, comparing the pre-2000 (royal blue) and post-2000 (orange) periods. Most fields exhibit substantial increases in the average proportion of causal edges over time. Fields such as Health, Urban, Development, and Behavioral show the largest increases and highest post-2000 levels. These patterns suggest that the adoption of causal inference methods has become more widespread across various fields in economic research, reflecting the broader impact of the credibility revolution.

This substantial increase suggests a growing emphasis on establishing credible causal relationships in economic research. The proliferation of empirical methods and a heightened focus on causal identification strategies have contributed to this trend, reflecting the impact of the credibility revolution in economics.

We also examine how this proportion varies across different fields within economics. Figure 3 Figure 3(b) presents the average proportion of causal edges by field, comparing the pre-2000 and post-2000 periods. Recent evidence corroborates these cross-field patterns, showing that finance and macroeconomics have lagged behind applied micro in adopting quasi-experimental methods, though the gap is narrowing ([36]). Our findings add that most fields have experienced substantial increases in the average proportion of causal edges in the post-2000 period.

Fields such as Urban, Health, Development, and Behavioral exhibit the highest increases. Urban increased from approximately 4.7% pre-2000 to 33.2% post-2000, marking one of the largest gains. Health saw an increase from 10.1% to 37.6%, achieving the highest post-2000 level among all fields. Development rose from 4.4% to 31.0%, and Behavioral increased from 3.6% to 29.6%. These fields, which often address policy-relevant questions and benefit from natural experiments or data conducive to causal analysis, have embraced causal inference methods more extensively.

Conversely, some fields experienced smaller increases or maintained lower levels. Macroeconomics increased modestly from 3.3% to 8.4%, reflecting a more cautious adoption of causal inference methods, possibly due to challenges in experimental design and identification strategies in macroeconomic contexts. Econometrics saw an increase from 6.3% to 11.0%, and Finance rose from 2.9% to 17.0%. These patterns suggest that the adoption of causal inference methods has varied across fields, influenced by the nature of the research questions, data availability, and methodological traditions within each field.

(Figure 3)

Evolution of Empirical Methods in Economic Research

To explore the increasing focus towards causal inference, we show time trends across methods and fields. Figure 4 illustrates the adoption of prominent empirical methods in NBER and CEPR working papers from 1980 to 2023. Methods such as DiD, IV, RCTs, and RDDs have seen substantial growth, reflecting the discipline’s shift towards more rigorous identification strategies.25

Proliferation of Empirical Methods Over Time in NBER and CEPR Working Papers

Figure 4. Note: This figure shows the proliferation of key empirical methods used in NBER and CEPR working papers over time: Difference-in-Differences (DiD), Instrumental Variables (IV), Randomized Controlled Trials (RCTs), Regression Discontinuity Design (RDD), Two-Way Fixed Effects (TWFE), Structural Estimation, Event Studies, Simulations, and Theoretical/Non-Empirical research. Each panel represents the proportion of papers utilizing one of these methods per year, with the y-axis showing the proportion of total papers and the x-axis indicating the year of publication. The data covers all NBER and CEPR working papers from 1980 to 2023. DiD has seen a significant increase since the 1980s, rising from around 4% to over 15% of papers in recent years, reflecting its growing importance in empirical research. IV methods have also increased steadily from approximately 2% to over 6% over the same period. RCTs and RDDs, while starting from near zero in the 1980s, have grown to over 7% and 2% respectively in recent years, indicating the rising feasibility and acceptance of experimental and quasi-experimental designs in economics. Conversely, the use of theoretical and non-empirical research has declined significantly, from around 20% in 1980 to under 10% in 2023, suggesting a shift towards empirical analysis in the discipline. The use of simulations has decreased from over 6% in 1980 to around 2–4% in recent years. These trends highlight the increasing emphasis on credible identification strategies and the evolution of empirical methods in economics.

DiD has become increasingly prevalent, rising from approximately 4% of papers in 1980 to over 15% in recent years. This growth underscores DiD’s utility in exploiting policy changes and natural experiments to identify causal effects. IV methods have also seen a steady increase, from around 2% of papers in 1980 to over 6% by 2023, highlighting their role in addressing endogeneity through exogenous instruments ([9]). The adoption of RCTs has accelerated since the early 2000s, increasing from less than 1% of papers in 2000 to over 7% in 2023, signaling the increasing feasibility and acceptance of experimental designs in economics. Similarly, RDD usage has grown from near zero in the 1980s to over 2% in recent years.

Conversely, the proportion of theoretical and non-empirical work has declined significantly, from approximately 20% of papers in 1980 to under 10% in 2023, indicating a broader emphasis on empirical analysis. The use of simulations has also decreased, from over 6% in 1980 to around 2–4% in recent years, possibly due to the growing availability of of data as well as the possibility to conduct (field) experiments for evaluation.

(Figure 4)

These trends are not uniform across subfields. Figure 5 presents the distribution of empirical methods by field. Fields such as Labour, Public, and Urban heavily utilize DiD, with over 12% of papers employing this method in Labour and Public, and over 16% in Urban. RCTs are particularly prominent in Behavioral and Development, where they are used in over 20% and 11% of papers respectively, reflecting the feasibility and policy relevance of experimental interventions in these areas. This trend aligns with the findings of [25], who documented a broad increase in empirical approaches due to the availability of big data and advanced computing.

Cross-sectional breakdown of empirical methods by field in NBER and CEPR working papers

Figure 5. Note: This figure displays the cross-sectional distribution of nine empirical methods—Difference-in-Differences (DiD), Instrumental Variables (IV), Randomized Controlled Trials (RCTs), Regression Discontinuity Design (RDD), Event Studies, Simulations, Structural Estimation, Two-Way Fixed Effects (TWFE), and Theoretical/Non-Empirical research—across twelve fields in NBER and CEPR working papers. Each point represents the proportion of papers within a specific field that utilize a given method, with 95% confidence intervals depicted by error bars. The fields include Finance, Development, Labour, Public, Urban, Macroeconomics, Behavioral, Economic History, Econometrics, IO, Environmental, and Health. The plot highlights considerable variation in the adoption of empirical methods across fields. DiD is most commonly used in Health, Urban, and Labour, with over 21% of papers in Health, over 16% in Urban, and over 13% in Labour utilizing this method. RCTs are particularly prominent in Behavioral and Development, where they are used in over 20% and 11% of papers respectively, reflecting the feasibility of experimental interventions in these areas. Simulations and Structural methods are more prevalent in Macroeconomics and Econometrics, reflecting the need for complex theoretical modeling in these fields. Simulations account for over 6% of papers in Macroeconomics and over 6% in Econometrics. Structural methods are used in approximately 6% of papers in Macroeconomics and over 5% in Econometrics. Fields like Macroeconomics and Finance rely more on IV methods and simulations, with Macroeconomics having around 3% of papers using IV methods and over 6% using simulations. Theoretical and non-empirical research remains significant in fields like Industrial Organization and Macroeconomics, with over 24% and 17% of papers respectively. These cross-sectional patterns reflect the methodological preferences specific to the research questions and data availability in each field, underscoring how different areas of economics adopt various empirical strategies to address their unique challenges.

In contrast, fields like Macroeconomics, IO, and Finance rely more on theoretical models and simulations to infer causal relationships from observational data. Theoretical approaches account for around 18% of papers in Macroeconomics and 12% in Finance. Structural estimation and simulations remain important in Macroeconomics and Industrial Organization, where complex theoretical models may be essential for understanding aggregate phenomena and market dynamics.

(Figure 5)

Paper-level Graphical Measures

In this section, we define quantitative measures that characterize the structure and content of each paper’s knowledge graph. These measures are grouped into three categories, each capturing distinct aspects of how research is organized and perceived by readers and editors. Measures of narrative complexity assess the internal structure of each paper’s claims; novelty and contribution compare a paper’s ideas to the existing literature, identifying new concepts or intersections; and conceptual importance and diversity situate a paper’s nodes within the broader knowledge graph of economic research.26 Together, these measures set the stage for the empirical analysis in Section 6, where we link them to publication outcomes and citation impacts. This section examines how these measures vary across fields and over time, presenting descriptive trends and cross-sectional comparisons to motivate their relevance for understanding academic reception.

Measures of Narratives Complexity

We consider several key measures derived from these claim networks. The number of edges, denoted as ∣Ep∣\lvert E_p\rvert, represents the total number of claims made in a paper, reflecting the breadth of the narrative. An increase in ∣Ep∣\lvert E_p\rvert suggests a more complex narrative with multiple interrelated claims. Similarly, the number of causal edges, denoted as ∣Epcausal∣\lvert E_p^{\mathrm{causal}}\rvert, represents the total number of claims evidenced using causal inference methods, indicating the depth of causal analysis within the paper.

Other measures include the number of unique paths in GpG_p, denoted as PpP_p, which is the total number of distinct directed paths between all pairs of nodes, excluding self-loops. This captures the interconnectedness of the narrative within the paper; a higher PpP_p indicates a more intertwined argument structure. The longest path length in GpG_p, denoted as LpL_p, represents the length of the longest directed path, indicating the depth of reasoning in the paper. We compute both PpP_p and LpL_p for the non-causal subgraph (denoted as Ppnon-causalP_p^{\mathrm{non\text{-}causal}} and Lpnon-causalL_p^{\mathrm{non\text{-}causal}}) and for the subgraph consisting only of causal edges (PpcausalP_p^{\mathrm{causal}} and LpcausalL_p^{\mathrm{causal}}).

Illustrative examples To concretely illustrate our graphical framework and the measures derived from it, we examine four landmark economic papers: [24], [13], [32], and [35]. These papers cover a diverse range of topics and methodologies, showcasing the versatility of our approach. The visual representations of these knowledge graphs are provided in Figure 6 and Figure 7, which highlight the varying structures and complexities of the narratives in these influential papers.

Knowledge graphs of two landmark economic papers, showing panels for Chetty et al. (2014) and Banerjee et al. (2015)

Figure 6. Knowledge Graphs of Two Landmark Economic Papers (Part 1). (a) [24] (2014). (b) [13] (2015).

Note: This figure presents the knowledge graphs of two landmark economic papers. Causal relationships are shown in orange, non-causal relationships in blue. Nodes represent economic concepts mapped to JEL codes; arrows indicate the direction of claims from source to sink. Panel (a) displays the graph for [24] (2014), showcasing multiple factors associated with upward mobility in the United States. The graph has ∣Ep∣=7|E_p| = 7 edges (all non-causal), Pp=6P_p = 6 unique paths, and a longest path length of Lp=1L_p = 1, indicating a broad but direct exploration of associations without extended causal chains. Panel (b) shows the graph for [13] (2015), illustrating the causal impact of introducing microfinance in India. The graph has ∣Ep∣=8|E_p| = 8 edges (all causal), Pp=12P_p = 12 unique paths, and a longest path length of Lp=3L_p = 3, reflecting a complex causal narrative with multiple interconnected outcomes resulting from the intervention.

Knowledge graphs of two landmark economic papers, with panel (a) Gabaix (2011) showing blue non-causal relationships among economic activities of large firms (D22), aggregate fluctuations (E10), idiosyncratic shocks do not average out (E32), fat-tailed distribution of firm sizes (D39), growth rate of GDP (O40), granular residual (L72), idiosyncratic shocks to large firms (E32), aggregate volatility (E10), and GDP fluctuations (F44); panel (b) Goldberg et al. (2010) showing orange causal and blue non-causal relationships among lower input tariffs (F13), increase in firms' product scope (L25), improved firm performance (output, TFP, R&D activities) (L25), declines in input tariffs (F14), increase in the scope of production by domestic firms (L25), increased availability of new imported varieties of inputs (O39), lower production costs (D24), and relax technological constraints for domestic firms (O25)

Figure 7. Knowledge Graphs of Two Landmark Economic Papers (Part 2). (a) [32]; (b) [35]. Note: This figure presents the knowledge graphs of two landmark economic papers. Causal relationships are shown in orange, non-causal relationships in blue. Nodes represent economic concepts mapped to JEL codes; arrows indicate the direction of claims from source to sink. Panel (a) presents the graph for [32], depicting theoretical relationships in macroeconomics concerning the impact of idiosyncratic firm-level shocks on aggregate economic fluctuations. The graph has ∣Ep∣=6|E_p| = 6 edges (all non-causal), Pp=11P_p = 11 unique paths, and a longest path length of Lp=3L_p = 3, indicating a complex theoretical narrative with deeper reasoning chains. Panel (b) displays the graph for [35], focusing on the effects of input tariff reductions on Indian firms’ product growth and performance. The graph has ∣Ep∣=5|E_p| = 5 edges (3 causal), Pp=5P_p = 5 unique paths, and a longest path length of Lp=2L_p = 2, reflecting a focused exploration of specific causal relationships supported by empirical methods.

:::

In [24], the authors investigate the geography of intergenerational mobility in the United States using administrative records. The paper presents a comprehensive analysis of how various factors correlate with upward mobility. The knowledge graph for this paper (fig:6a) includes seven edges, all of which are non-causal. The relationships mapped include how parent income (JEL code D31) influences child income rank (J13), and how factors such as lower residential segregation (R23), less income inequality (D31), better primary schools (I21), greater social capital (Z13), and more stable family structures (J12) are associated with higher upward mobility (J62). The knowledge graph measures for this paper are as follows: the number of edges ∣Ep∣=7|E_p| = 7 (all non-causal), the number of unique paths Ppnon−causal=6P_p^{\mathrm{non-causal}} = 6, and the longest path length Lpnon−causal=1L_p^{\mathrm{non-causal}} = 1. The relatively high number of edges indicates a broad exploration of factors affecting upward mobility, while the longest path length of 1 reflects that the relationships are primarily direct associations rather than extended causal chains.

[13] report on a randomized evaluation of the impact of introducing microfinance in a new market. Utilizing Randomized Controlled Trials (RCTs), the authors assess how access to microfinance affects various economic outcomes for households in Hyderabad, India. The knowledge graph for this paper (fig:6b) includes eight edges, all of which are causal, evidenced through RCTs. The causal relationships include how the introduction of microfinance (G21) leads to households having a microcredit loan (D14), which in turn influences new business creation (L26) and investment in existing businesses (G31). These investments impact average monthly per capita expenditure (D12), while microcredit affects expenditure on durable goods (E21) and expenditure on temptation goods (D12). The authors also examine the effect of microcredit on development outcomes such as health, education, and women’s empowerment (I15). The knowledge graph measures are: the number of edges ∣Ep∣=8|E_p| = 8 (all causal), the number of unique paths Ppcausal=12P_p^{\mathrm{causal}} = 12, and the longest path length Lpcausal=3L_p^{\mathrm{causal}} = 3. The high number of causal edges and unique paths indicates a complex causal narrative with multiple interconnected outcomes, while the longest path length reflects deeper causal chains.

(Figure 6)

In [32] (2011), the author proposes that idiosyncratic firm-level fluctuations can explain a significant part of aggregate economic shocks, introducing the "granular" hypothesis. The paper develops a theoretical framework and provides empirical evidence supporting the idea that shocks to large firms contribute to aggregate volatility. The knowledge graph (fig:7a) includes six edges, all of which are non-causal, representing theoretical relationships such as how idiosyncratic shocks to large firms (D21) contribute to aggregate volatility (E32) and GDP fluctuations (F44), and how the fat-tailed distribution of firm sizes (L11) implies that idiosyncratic shocks do not average out (E32). The knowledge graph measures are: the number of edges ∣Ep∣=6|E_p| = 6 (all non-causal), the number of unique paths Ppnon−causal=11P_p^{\mathrm{non-causal}} = 11, and the longest path length Lpnon−causal=3L_p^{\mathrm{non-causal}} = 3. The relatively high number of unique paths and the longest path length indicate a complex theoretical narrative with multiple interconnected concepts and deeper reasoning chains.

Finally, [35] (2010) examine the impact of imported intermediate inputs on domestic product growth in India. The authors use empirical methods, including Difference-in-Differences (DiD), to establish causal relationships between declines in input tariffs and firm performance. The knowledge graph (fig:7b) for this paper includes five edges, three of which are causal. The causal relationships include how declines in input tariffs (F14) lead to increased firm product scope (L25) and improved firm performance (L25), and how increased availability of new imported inputs (O39) causes relaxed technological constraints for domestic firms (O33). The knowledge graph measures are: the number of edges ∣Ep∣=5|E_p| = 5, the number of causal edges ∣Epcausal∣=3|E_p^{\mathrm{causal}}| = 3, the number of unique paths Pp=5P_p = 5, and the longest path length Lp=2L_p = 2. These measures reflect a focused exploration of specific causal relationships, with a moderate level of narrative complexity.

(Figure 7)

These examples demonstrate the diversity in the structure and complexity of knowledge graphs across different types of economic research. They illustrate how our measures capture key aspects of the narratives, such as the breadth of topics covered, the depth of causal analysis, and the interconnectedness of concepts.

Measures of Novelty and Contribution to Literature

Thus far, our measures of narrative complexity focus on the internal structure of each paper’s knowledge graph, such as the number of edges, unique paths, and longest causal chains. We now introduce a complementary set of measures capturing how each paper pushes the frontier of economic research by contributing new concepts, links, or underexplored intersections. These measures fall into four categories: novel edges, path-based novelty, subgraph-based novelty, and gap filling. Each can be computed for the non-causal subgraph or restricted to the causal subgraph. In the latter case, these measures capture novelty or gap filling specifically within the subset of claims that are supported by causal inference methods.

Novel Edges

A paper’s knowledge graph may contain numerous directed edges (claims). We call an edge (u→v)(u \to v) novel if it was never previously documented in any earlier paper’s knowledge graph. For instance, if no prior publication had directly linked the concept of Climate Change to Migration Flows, the first paper to do so would be credited with introducing a novel edge.

Formally, let ⋃q<pEq\bigcup_{q<p} E_q be the union of all edges from papers preceding pp. Then an edge (u→v)∈Ep(u \to v) \in E_p is novel if (u→v)∉⋃q<pEq(u \to v) \notin\bigcup_{q<p} E_q.

We define two statistics:

NumNovelEdges⁡p=#{(u→v)∈Ep:(u→v) is novel},\operatorname{NumNovelEdges}_{p} = \#\{(u \to v) \in E_p : (u \to v)\text{ is novel}\},
PropNovelEdges⁡p=NumNovelEdges⁡p∣Ep∣.\operatorname{PropNovelEdges}_{p} = \frac{\operatorname{NumNovelEdges}_{p}}{\lvert E_p\rvert}.

PropNovelEdgespPropNovelEdges_p thus captures the fraction of a paper’s edges that represent wholly new linkages between concepts, relative to the existing corpus.

Algorithmically, we maintain a global (yearly) list of edges seen so far. For each new paper pp, we check each edge against that list to see if it is new. We then record the proportion of edges that are novel. We implement an identical procedure for both the causal and non-causal subgraph by filtering edges to those based on causal inference methods.

Suppose prior literature had already established claims linking Income Inequality and Social Capital, but no one had linked Family Stability directly to Upward Mobility. If [24] were the first to draw that edge, it counts as a novel linkage, thus incrementing NumNovelEdges.

Path-Based Novelty

Even if the individual edges of a paper are not entirely new, the paper may still combine known edges into new sequences of concepts that have never before appeared in a single chain. For instance, two edges (X→Y)(X \to Y) and (Y→Z)(Y \to Z) might be known, but if no prior paper has ever connected them into the path (X→Y→Z)(X \to Y \to Z), that path is new. This captures a higher-order notion of novelty: novel paths.

For each paper pp, let Pp\mathcal{P}_p be all simple directed paths in GpG_p up to a chosen length ℓ\ell. A path π∈Pp\pi\in\mathcal{P}_p is novel if it did not appear in ⋃q<pPq\bigcup_{q<p}\mathcal{P}_q. We then define:

NumNovelPaths⁡p=∣{π∈Pp:π is novel}∣andPropNovelPaths⁡p=NumNovelPaths⁡p∣Pp∣\operatorname{NumNovelPaths}_p = \left|\{\pi\in\mathcal{P}_p : \pi\text{ is novel}\}\right| \quad\text{and} \quad\operatorname{PropNovelPaths}_p = \frac{\operatorname{NumNovelPaths}_p}{|\mathcal{P}_p|}

In terms of implementation, for each paper we first build the directed graph and list all simple paths up to length ℓ\ell (often ℓ=3\ell= 3). Secondly, we convert each path into a canonical string representation (e.g., “u→v→wu \to v \to w”), and compare these strings to a global set of previously recorded paths. Again, we do this for both the non-causal subgraph and the causal subgraph.

This measure helps us think about novelty in joint mechanisms. [67] famously argued that scientific ideas grow combinatorially as researchers recombine existing concepts in novel ways, resonating with our notion that bridging seldom-linked nodes fosters innovation. Prior work might have had edges from Microfinance Introduction to Borrowing Behavior, and from Borrowing Behavior to Business Investment. But if no paper had combined these into a causal chain

Microfinance →\to Borrowing →\to Business Investment, then [13] would score highly on path-based novelty.

Subgraph-Based Novelty

Subgraphs capture multi-node patterns that may include loops or branching structures. A subgraph S=(VS,ES)S = (V_S, E_S) of GpG_p is said to be novel if no isomorphic subgraph has appeared in prior papers. For computational reasons, we typically limit attention to subgraphs with ∣VS∣≤k|V_S| \le k (e.g., 3 or 4 nodes). Subgraph enumeration and motif detection have long been central to the study of complex networks in fields ranging from systems biology to social network analysis.27

Let SpS_p be the set of all induced subgraphs of size kk within GpG_p. We produce a canonical representation ϕ(S)\phi(S) for each subgraph to check for duplicates. We then define:

NovelSubgraphs⁡p={ϕ(S)∈Sp:ϕ(S)∉Φ<p},PropNovelSubgraphs⁡p=∣NovelSubgraphs⁡p∣∣Sp∣\operatorname{NovelSubgraphs}_p = \left\{\phi(S) \in S_p : \phi(S) \notin\Phi_{<p}\right\}, \qquad\operatorname{PropNovelSubgraphs}_p = \frac{\left|\operatorname{NovelSubgraphs}_p\right|}{\left|S_p\right|}

where Φ<p\Phi_{<p} is the union of canonical subgraph signatures from all earlier papers.

The algorithm enumerates all subgraphs of size kk in each paper using an induced_subgraph procedure and checks for isomorphism by generating a canonical adjacency-matrix signature. We then cumulatively union these signatures across papers in chronological order. As before, we produce separate measures for the non-causal subgraph and for the causal subgraph.

Unlike individual edges or single paths, subgraphs can capture intricate feedback loops or multi-node structures. For instance, a triad of nodes

Financial Constraint→Firm Innovation→Economic Growth→Financial Constraint\mathit{Financial\ Constraint} \to\mathit{Firm\ Innovation} \to\mathit{Economic\ Growth} \to\mathit{Financial\ Constraint}

would be a 3-node cyclic subgraph. If no prior paper documented this three-way feedback cycle, the paper introducing it obtains a high subgraph-based novelty measure.

Gap Filling (Co-occurrence Analysis)

Whereas the above novelty measures track new edges or structures, gap filling identifies whether a paper bridges concepts that are underexplored in prior work. Even if each concept is well studied on its own, a pair might be rare or entirely missing in the earlier literature. A paper that explicitly connects these underexplored concepts helps fill that gap. In economics, the use of co-occurrence methods is prominently exemplified by the work of [39] in measuring export complexity, but these approaches are more widely applied in fields such as biomedical research, bibliometrics, and network science.28

For each year, we compile a cumulative count of how many times each unordered pair of JEL codes has co-occurred in prior papers. We label any pair {u,v}\{u,v\} with frequency below a threshold τ\tau (e.g., fewer than 5 co-occurrences) as underexplored. Given a new paper pp, we then look at all distinct pairs of JEL codes in VpV_p and calculate the gap-filling proportion as the fraction of these pairs that remain underexplored:

GapFillingProp⁡p=#{{u,v}∈Vp2:count⁡(u,v)<τ}#{{u,v}∈Vp2}.\operatorname{GapFillingProp}_{p} = \frac{\#\{\{u,v\} \in V_p^2 : \operatorname{count}(u,v) < \tau\}}{\#\{\{u,v\} \in V_p^2\}}.

A higher value indicates that the paper bridges more underexplored concept pairs.

Algorithmically, we maintain an updated frequency table for all unordered pairs of JEL codes up to year t−1t-1. For each paper published in year tt, we check which pairs in its node set remain below the threshold τ\tau and compute that paper’s gap-filling proportion. The same procedure is applied to the causal subgraph if one wishes to isolate underexplored pairs specifically tied to credible-causal edges.

Consider the following example. A macroeconomics paper linking Behavioral Biases (D91) and Monetary Policy (E52) in a new structural model might fill a gap if these two codes historically co-occur rarely (e.g., fewer than 5 times). Even if prior literature studied each code extensively in isolation, the explicit combination might represent a novel domain intersection.

Which Journals Fill Gaps in Literature? To further illustrate our gap-filling measure, we compare how papers accepted in different journal tiers (and specifically among the top 5 journals) bridge underexplored connections. Figures fig:8a and fig:8b display results for gap filling (Section 5.2.4). In each panel, we report separate estimates for the non-causal subgraph (in orange) and the causal subgraph (in blue). The error bars represent 95% confidence intervals.29

In Panel fig:8a, the Top 5 group shows a higher mean of standard gap filling for the non-causal subgraph, but we observe very little difference across tiers for the causal subgraph. This suggests that top-tier outlets more readily reward bridging underexplored concept pairs in general, although the strictly causal bridging does not differ much between tiers. Panel fig:8b disaggregates the five elite journals, highlighting especially strong causal gap filling in the QJE and AER.

Overall, these patterns suggest that top-tier outlets reward papers forging novel or underexplored concept connections, especially when supported by credible identification strategies.

(Figure 8)

Gap Filling Across Journal Quality Tiers and Within Top 5 Journals

Figure 8. Note: Each panel displays average gap-filling proportions (%) with 95% confidence intervals, distinguishing between non-causal edges (orange) and causal edges (blue). Panel (a) shows gap-filling proportions (co-occurrence based) across three journal tiers (Top 5, Top 6–20, Top 21–100). We find that Top 5 journals, on average, have higher non-causal and slightly higher causal gap-filling proportions. Panel (b) restricts attention to the top five outlets, showing that the QJE (particularly) and AER accept papers with notably higher gap-filling in the causal subgraph, compared to the other top journals. The causal gap filling proportion is more than three times greater at QJE (around 8%) than at Econometrica (around 2.5%).

Taken together, these measures of novelty and gap filling enrich our understanding of how each paper pushes the frontier of economic knowledge. By systematically tracking new edges, new paths, new subgraphs, and underexplored concept pairs, we capture distinct facets of a paper’s originality and the extent to which it addresses previously overlooked topics or mechanisms. As shown in Section 6, these indicators prove to be significant predictors of both publication outcomes and the long-run citation impact of research.

Measures of Conceptual Importance and Diversity

We now turn to assessing how centrally each paper’s concepts lie within the broader economic literature and whether the paper balances multiple sources and targets in its argumentation. These attributes may shape both the paper’s perceived significance and its eventual scholarly influence. We consider two sets of measures: (i) the source–sink ratio, which captures the balance of causal flows in the paper, and (ii) centrality-based indicators (e.g., eigenvector or PageRank scores) computed for the nodes in a cumulative knowledge graph of prior literature. We then aggregate these centrality values (mean, variance) over the specific concepts a paper employs, providing measures of conceptual importance and conceptual diversity.

Source–Sink Ratio

A paper may describe multiple causal mechanisms converging on a single outcome or, conversely, a narrow set of causes that generate multiple distinct outcomes. We define a source–sink ratio to capture these tendencies. Let Gp=(Vp,Ep)G_p = (V_p, E_p) be the knowledge graph of paper pp. For each node v∈Vpv \in V_p, let outdegree⁡(v)\operatorname{outdegree}(v) and indegree⁡(v)\operatorname{indegree}(v) denote its out-degree and in-degree, respectively. The source–sink ratio is then:

Rp=∣{v∈Vp:outdegree⁡(v)>0}∣∣{v∈Vp:indegree⁡(v)>0}∣+εR_p = \frac{\left|\left\{v \in V_p : \operatorname{outdegree}(v) > 0\right\}\right|}{\left|\left\{v \in V_p : \operatorname{indegree}(v) > 0\right\}\right| + \varepsilon}

where ε\varepsilon is a small positive constant to avoid division by zero if no nodes act as sinks (or sources).

A ratio Rp>1R_p > 1 indicates that the paper has more distinct source nodes (origins of claims) than sink nodes (endpoints). This often arises in work that investigates many potential drivers of a smaller set of outcomes. Conversely, Rp<1R_p < 1 suggests the paper concentrates on a few key causal factors but explores a broad set of consequences. A balanced ratio around 1.0 reflects a more even structure, with multiple causes and effects.

We compute RpR_p separately for the non-causal subgraph and the causal subgraph. The latter includes only edges established via credible identification strategies (e.g., DiD, IV, or RDD). A paper that juggles multiple rigorously identified drivers for a single main outcome, for instance, will display a high causal source–sink ratio.

If a paper focuses heavily on how multiple institutional features (e.g., tax incentives, labor regulations, trade reforms) affect a single macroeconomic indicator (like GDP growth), it may have a high source–sink ratio. By contrast, if it starts from one main policy shock but explores myriad downstream effects (household consumption, firm investment, labor demand, etc.), the ratio can be well below 1.0.

To examine how the source–sink ratio evolves, Figure 9 plots its average values (along with 95% confidence bands) for the non-causal and causal subgraphs over time. The ratio for the non-causal subgraph remains relatively flat until 2000, followed by a slight decline in subsequent years. In contrast, the causal subgraph shows a steady and steep increase over the same period, indicating that recent decades see papers exploring more distinct causal factors relative to outcomes. This pattern aligns with a shift toward identifying multiple potential mechanisms or drivers for a narrower set of effects.

Image not extracted

Figure 9.

Figure 9 compares fields cross-sectionally. We see that Econ History feature high non-causal source-sink ratio, given its emphasis on explaining a single outcome through multiple correlational determinants. By contrast, Behavioral and Health exhibit larger causal source-sink ratios, reflecting the broader adoption of rigorous identification that tests several potential causal channels for specific outcomes.30 Together, these field-level differences highlight how some areas of economics appear more conducive to multi-causal designs, while others continue to aggregate a range of correlates into narrower outcome measures.

Centrality-Based Measures of Conceptual Importance and Diversity

Network centrality indices assess how “important” or “influential” a node is within a larger network ([17]; [52]). In this context, a node is a JEL code (economic concept), and the larger network is the cumulative knowledge graph of all prior economic research in our sample. We compute centralities for each concept in the network at year t−1t-1, then measure how central (on average) the concepts used by paper pp (published in year tt) actually are. A higher average centrality indicates that the paper engages with well-established or highly influential concepts, whereas a lower average centrality may signal an exploration of more specialized or peripheral topics. Furthermore, a higher variance of centralities among the paper’s nodes can imply mixing mainstream and niche concepts—a form of conceptual diversity.

We focus primarily on eigenvector centrality because it captures the recursive intuition that a node is important if it is connected to other important nodes. However, we also compute PageRank [52] to check robustness, since PageRank includes a damping factor that prevents rank sinks and redistributes influence among nodes with low in-degree. In directed citation or claim networks, PageRank can better handle large cycles and “dangling” nodes, but both measures reflect similar underlying principles of node influence.

Formally, let Ct−1(v)C_{t-1}(v) be the centrality of node vv in the cumulative knowledge graph formed by all papers prior to year tt. Once we have identified the set of JEL codes VpV_p used by paper pp, we define:

MeanCentrality⁡p=1∣Vp∣∑v∈VpCt−1(v),VarCentrality⁡p=1∣Vp∣∑v∈Vp[Ct−1(v)−MeanCentrality⁡p]2.\operatorname{MeanCentrality}_p = \frac{1}{\lvert V_p\rvert}\sum_{v\in V_p} C_{t-1}(v), \qquad\operatorname{VarCentrality}_p = \frac{1}{\lvert V_p\rvert}\sum_{v\in V_p} [C_{t-1}(v)-\operatorname{MeanCentrality}_p]^2.

These measures are computed for two types of subgraphs: the non-causal subgraph and the causal subgraph. Specifically, we compute MeanCentrality⁡p\operatorname{MeanCentrality}_p and VarCentrality⁡p\operatorname{VarCentrality}_p using eigenvector centrality and PageRank for both the non-causal and causal subgraphs. This yields the following statistics:

MeanEigen⁡p,VarEigen⁡p,MeanPR⁡p,VarPR⁡p,\operatorname{MeanEigen}_p,\quad\operatorname{VarEigen}_p,\quad\operatorname{MeanPR}_p,\quad\operatorname{VarPR}_p,

calculated separately for the non-causal and causal subgraphs. By implementing this procedure for each subgraph type, we capture differences in conceptual importance and diversity for claims supported by causal inference methods versus those based on non-causal relationships.

Once we calculate a paper’s mean or variance of eigenvector (or PageRank) centralities, we standardize these values (i.e., subtract the sample mean and divide by the standard deviation) to facilitate comparability across papers and subgraphs.

Nodes with high prior-year centrality are typically well-established topics (e.g., GDP, Unemployment, Trade Liberalization, Exchange Rates). Papers engaging predominantly with such nodes may attract broader readership, facilitate connections to familiar debates, and potentially garner more citations. By contrast, papers emphasizing less central nodes (e.g., Renewable Energy Technologies in 1990, or Machine Learning for Macroeconomic Forecasting in 2010) might be exploring frontier topics with sparser prior literature. The variance measure captures whether a paper blends mainstream concepts with more specialized ones. This diversity may be an asset (appealing to multiple subfields) or a challenge (positioning the paper less neatly in any single mainstream area).

Central concepts can serve as natural anchors in scientific discourse, accruing more attention and building cumulative knowledge ([49]). Engaging with such mainstream nodes may boost impact, but may also reduce a paper’s perceived novelty. In turn, venturing into peripheral areas can differentiate one’s research but risks a narrower audience. Our empirical analysis in Section 6 demonstrates that these centrality-based indicators correlate systematically with journal placement and citation trajectories, highlighting how conceptual positioning can shape both short-run visibility and long-run influence.

Illustrative examples

To demonstrate how these measures vary across papers, we revisit the four landmark studies while focusing on the source–sink ratio, average eigenvector centrality, and the variance of centrality within nodes. In [24], the source-sink ratio Rp≈2.5R_p \approx2.5 is derived from the non-causal subgraph, as this paper contains only non-causal edges. This indicates multiple sources—e.g., residential segregation, income inequality, school quality, social capital, family stability—culminating in a single main outcome (upward mobility). Their standardized average eigenvector centrality is about −0.39-0.39, with a variance of around −0.46-0.46.31 Although below the global mean, this suggests moderate reliance on somewhat less central concepts in the broader non-causal knowledge graph (e.g., intergenerational mobility, which, at the time, was a relatively specialized research topic).

By contrast, [13] exhibit a causal source–sink ratio of Rp≈0.71R_p \approx0.71, reflecting a single main cause (microfinance introduction) that branches out into multiple downstream outcomes, such as borrower behavior, business creation, expenditure patterns, and development indicators. Their standardized average eigenvector centrality is approximately −0.35-0.35 (variance: −0.22-0.22), placing these concepts slightly below the global causal mean. This modestly negative value aligns with the paper’s relatively frontier policy topics—e.g., finance for the poor—which, although influential, had not yet reached the level of established mainstream concepts.

In [32], Rp≈0.8R_p \approx0.8 (non-causal) indicates a mix of firm-level shocks feeding into macro-level outcomes, such as aggregate volatility and GDP fluctuations. The standardized average eigenvector centrality is about −0.15-0.15 (variance: −0.14-0.14), reflecting more specialized or less frequently studied areas (e.g., the fat-tailed distribution of firm sizes) prior to that paper’s publication. Although these concepts ultimately sparked considerable interest, they were comparatively less central within the broader non-causal network at the time.

Finally, [35] highlight multiple sources—declines in input tariffs and increased availability of new imported inputs—affecting a narrower set of firm outcomes (product scope, performance). For the non-causal subgraph, the source–sink ratio is around 1.33, with a standardized mean eigenvector centrality of about −0.60-0.60 (variance: −0.26-0.26). Meanwhile, the causal subgraph yields a higher source–sink ratio of roughly 2.0, alongside a standardized mean eigenvector centrality of −0.57-0.57 (variance: −0.24-0.24). Although both measures sit below the overall mean, the concepts involved (e.g., trade liberalization, industrial performance) do have moderate footing in established debates, consistent with the paper’s balanced mix of theoretical and empirical contributions.

These examples illustrate how structural measures of source–sink balance and centralities capture different dimensions of a paper’s conceptual organization. By quantifying these properties, we gain a clearer picture of how particular studies draw on established conversations or venture into more specialized territories, and how such positioning may influence their academic reception.

Evolution of Concept Centrality Over Time in Literature

We next investigate how the centrality of specific economic concepts has changed over time, comparing both the non-causal and the causal subgraph. First, we highlight which JEL codes rank most highly in eigenvector centrality, noting differences between non-causal edges and those established through causal inference methods. We then track how certain concepts have risen or fallen in importance since 1990.

Most central concepts in economics Figure 10 displays the top 20 JEL codes ranked by their eigenvector centrality in both the non-causal and causal knowledge graphs of the literature (i.e., our full sample, partitioned by edge type). The comparison reveals interesting patterns in how central economic concepts differ based on whether claims are supported by causal inference.

Top 20 JEL Codes by Eigenvector Centrality in Overall and Causal Knowledge Graphs

Figure 10. Top 20 JEL Codes by Eigenvector Centrality in Overall and Causal Knowledge Graphs. Note: This figure displays the top 20 JEL codes ranked by their eigenvector centrality within both the non-causal knowledge graph (orange points) and the causal knowledge graph (blue points) constructed from our dataset. Eigenvector centrality measures the influence of a node in a network, with higher scores indicating more central or connected concepts. The comparison reveals that while nodes like G21 (Banks and Mortgages) and J31 (Wage Structure) have high centrality in the non-causal graph, topics such as I24 (Education and Inequality), J13 (Fertility and Family), and I21 (Analysis of Education) are more central in the causal graph. This suggests a shift in focus towards these areas when using causal inference methods in economic research.

In the non-causal subgraph, the nodes with the highest eigenvector centrality scores include G21 (Banks and Mortgages), J31 (Wage Structure), and I24 (Education and Inequality), reflecting their prominence in the economic literature at large. However, when focusing on the causal subgraph, the nodes with the highest centrality scores shift towards I24 (Education and Inequality), J13 (Fertility and Family), and I21 (Analysis of Education).

This shift indicates that while certain economic concepts are central when considering non-causal associations, the focus of causal inference methods tends to concentrate on specific areas, such as education, family, and health economics—topics more amenable to experimental or quasi-experimental designs.

The distribution of eigenvector centralities is highly skewed in both graphs. In the non-causal subgraph, the mean eigenvector centrality is 0.0402, with a maximum of 0.898 for G21. In the causal subgraph, the centralities are more concentrated among certain nodes, with I24 achieving the highest score normalized to 1.0. By incorporating average eigenvector centrality as a measure, we capture the extent to which a paper’s claims involve central or peripheral concepts in economics. Papers engaging with highly central nodes in the causal subgraph may be contributing to well-established areas of research using rigorous methods, while those involving less central nodes might be exploring more novel or specialized topics. (Figure 10)

Rise and Fall of Importance of Certain topics Beyond identifying which nodes are central at a given moment, we examine how these centralities evolve over time. Figure 11 plots time series for five nodes that rose most steeply in normalized eigenvector centrality from 1990 to 2020 and five that declined the most, based on changes between a baseline year (e.g., 2000) and the final year (2020). We perform a min–max normalization for each node to facilitate cross-node comparison on a 0–1 scale.

Top 5 Rising and Declining Nodes’ Eigenvector Centrality Over Time (Normalized)

Figure 11. Top 5 Rising and Declining Nodes’ Eigenvector Centrality Over Time (Normalized). Note: Each line represents one JEL code’s yearly eigenvector centrality, min–max normalized to highlight relative gains and losses between 1990 and 2020. Panel (a) shows results for the Non-Causal subgraph, where macro-oriented categories such as Growth & Productivity (O49) and Interest Rates (E43) show marked declines, while micro-oriented topics (Micro-based Behavioural Economics (D91), Minorities and Non-Labor Discrimination (J15), Health & Inequality (I14), etc.) trend upward. Panel (b) focuses on the Causal subgraph, restricted to edges supported by causal inference methods, and reveals a similar shift from traditional macro/finance topics toward areas linking economic outcomes to social, behavioral, and inequality themes.

Panel (a) shows results for the Non-Causal subgraph. Concepts related to traditional macro-economic or “hard” economic areas—such as Growth & Productivity (O49), Interest Rates (E43), Capital & Investment (E22), Foreign Exchange (F31), and Wages & Labor (J31)—all show declining centrality. By contrast, topics often associated with microeconomics, inequality, or behavioral dimensions—Health & Inequality (I14), Behaviour & Decisions (D91), Collective Decisions (D79), Minorities & Non-Labor Discrimination (J15), and Education & Inequality (I24)—exhibit sustained rises. This pattern suggests a structural shift away from classical macro–finance focal points toward more applied or “softer” fields linking economics to questions of health, social behavior, and inequality.

Panel (b) focuses on the Causal subgraph, restricted to edges supported by causal inference methods (e.g., DiD, IV, or RDD). The trends observed here are broadly similar, with declines in traditional macro/finance topics and rising importance for micro-oriented, behavioral, and inequality-focused areas. Notably, Health & Inequality (I14), Education & Inequality (I24), and related codes exhibit consistent increases in centrality, underscoring the growing emphasis on empirically identified relationships in policy-relevant areas. This shift reflects the broader influence of the credibility revolution in steering economic research toward topics where causal inference methods are particularly applicable.

(Figure 11)

Knowledge Graph Predictors of Publication and Citations

This section investigates how the structural properties of a paper’s knowledge graph predict its subsequent publication outcomes and citation impacts. Drawing on the measures introduced in Section 5, we examine three categories of predictors: narrative complexity, novelty and contribution, and conceptual importance and diversity. These categories capture distinct dimensions of how research is structured and how it connects to existing literature. We evaluate whether these features systematically correlate with a paper’s likelihood of publication in prestigious journals and its subsequent citation impact.

Empirical Framework and Variables

This section outlines the empirical strategy to evaluate how paper-level knowledge-graph measures predict publication outcomes and citation impacts. For each measure introduced in Section 5 (e.g., narrative complexity, novelty, conceptual importance), we separately analyze its non-causal and causal variants. The regression specification is:

yp=α+βMp(type)+δt(p)+εp,(1)y_p = \alpha+ \beta M_p^{(\mathrm{type})} + \delta_{t(p)} + \varepsilon_p, \tag*{(1)}

where ypy_p is one of four outcomes: (i) an indicator for publication in a top 5 journal (Top5pTop5_p); (ii) publication in a journal ranked 6–20 (Top6Top6-20p20_p); (iii) publication in a journal ranked 21–100 (Top21Top21-100p100_p); or (iv) the log of total citations plus one (log⁡(Citesp+1)\log(Cites_p + 1)). The explanatory variable Mp(type)M_p^{(\mathrm{type})} represents a paper-level measure (e.g., number of edges, proportion of novel paths, or average centrality), analyzed separately for non-causal and causal subgraphs.

We estimate two versions of this regression. The first includes year fixed effects, δt(p)\delta_{t(p)}, to control for time-specific shocks and trends in publication or citation practices, with standard errors clustered at the year level to account for within-year correlations among papers. The second version excludes year fixed effects and reports unclustered standard errors, allowing us to assess the robustness of our results to the inclusion of year-level controls. Both versions are included in the results for comparison.

This framework facilitates a detailed comparison of the predictive power of non-causal versus causal claims for publication and citation outcomes. While causal relationships often have greater influence on top-tier journal placement, both non-causal and causal measures significantly affect long-term citation impacts, as explored in the following subsections.

Narrative Complexity

Figure 12 presents the coefficient estimates for narrative complexity measures, which include the number of edges, the number of unique paths, and the log of the longest path length, analyzed separately for the non-causal and causal subgraphs. The results highlight that causal narrative complexity is positively associated with most publication and citation outcomes across both specifications (with and without year fixed effects), with particularly large effect sizes for publication in top 5 journals and citations. Among the three measures, the log of the longest path length has the largest effect size. For instance, with year fixed effects, a one-unit increase in the log of the longest path length is associated with around 0.05 increase in the probability of publication in a top 5 journal, which represents more than a fifth of the mean for this outcome (0.22). Similarly, the effect size for citations is substantial, with causal measures showing consistent positive and statistically significant coefficients.

Coefficient estimates relating narrative complexity measures to publication outcomes and citation counts

Figure 12. Relationship between Narrative Complexity Measures with Publication and Citations

Note: This figure shows coefficient estimates relating narrative complexity measures (number of edges, number of unique paths, longest path length) to publication outcomes (Top 5, Top 6–20, Top 21–100 journals) and citation counts. Estimates are displayed for both “Non-Causal” and “Causal” versions of the measures, and specifications with and without year fixed effects. Positive coefficients indicate that higher complexity is associated with higher likelihood of top-tier publication or greater citations. We find that causal narrative complexity is a key driver of both initial placement in top-tier journals and longer-term citation impact, with particularly large effects for the log of the longest path length. Non-causal narrative complexity, by contrast, exhibits no positive influence and, in some cases, is negatively correlated with outcomes.

In contrast, non-causal narrative complexity measures exhibit zero to negative coefficients across all outcomes. This suggests that general complexity unsupported by causal inference methods may not enhance publication or citation outcomes and, in some cases, may even hinder them.

Interestingly, the relationship with top 21–100 publications does not display the same consistent positive pattern. For that tier, causal narrative complexity measures have positive but small coefficients that are often statistically insignificant, particularly when year fixed effects are included. Non-causal measures remain consistently at or near zero, further indicating that lower-tier journals do not systematically reward narrative complexity. These results suggest that while complexity embedded in rigorously identified causal claims can improve placement in top-tier journals, its impact on mid-tier and field-specific outlets is limited.

Turning to citations, all causal narrative complexity measures are positively associated with citation counts, with large and statistically significant coefficients. This indicates that complexity grounded in credible causal relationships enhances longer-run academic influence. In contrast, non-causal measures are negatively associated with citations, particularly in the specification with year fixed effects, and are statistically insignificant without them. For example, the log of the longest path length in the causal subgraph has a substantial effect size on citations, demonstrating the importance of interconnected causal reasoning in generating scholarly attention.

(Figure 12)

Novelty and Gap Filling

Figure 13 presents the results for measures capturing novelty and gap filling. For novelty, we examine three measures: the proportion of new edges, the proportion of novel unique paths, and the proportion of novel subgraphs introduced by a paper. Consistently, causal novelty measures are positively associated with publication in top 5 journals, with large and statistically significant effect sizes, except for novel subgraphs, which exhibit a positive but statistically insignificant relationship. In contrast, non-causal novelty measures for top 5 journals are generally positive but often statistically indistinguishable from zero, highlighting the limited value placed on novelty unsupported by causal evidence.

Relationship between Novelty and Gap Filling Measures with Publication and Citations

Figure 13. Relationship between Novelty and Gap Filling Measures with Publication and Citations

Note: This figure displays coefficient estimates relating measures of novelty (e.g., proportion of new edges, novel subgraphs, novel paths) and gap filling to publication outcomes and citation counts. Results are shown for both “Non-Causal” and “Causal” measures and include specifications with and without year fixed effects. Positive coefficients indicate a positive association with top-tier publication or citations.

For top 6–20 journals, the results differ. Non-causal novelty measures are consistently negative, suggesting that such novelty may detract from placement in mid-tier journals. For causal novelty, the results are mixed. Novel subgraphs show negative but statistically insignificant coefficients, while novel edges and paths exhibit positive but small effect sizes. These patterns suggest that top 5 journals are more likely to reward novelty supported by causal inference methods, while mid-tier outlets exhibit weaker preferences for novelty, particularly for causal subgraph-based novelty.

When the outcome is publication in top 21–100 journals, the results are statistically insignificant and mixed across both non-causal and causal measures. This further reinforces the idea that lower-tier journals do not systematically reward novelty, regardless of whether it is causally grounded.

Turning to citations, a stark and consistent difference emerges between causal and non-causal novelty measures. Causal novelty measures are all positive and statistically significant, with large effect sizes, demonstrating the importance of introducing new causal relationships in generating long-term scholarly attention. By contrast, non-causal novelty measures are uniformly negative, underscoring the limited influence of general novelty unsupported by causal identification on a paper’s academic impact.

The gap filling measure examines whether papers bridge previously disconnected concepts. Results show that top 5 journals reward both causal and non-causal gap filling, with positive and statistically significant coefficients for both types. However, for top 6–20 and top 21–100 journals, only the causal gap filling measure retains a positive or zero-to-positive relationship, while the non-causal measure is zero or negative. Interestingly, neither causal nor non-causal gap filling measures are significantly associated with citation outcomes, suggesting that bridging underexplored concepts does not consistently translate into long-term academic influence.

Overall, the results indicate that novelty is a critical factor for top 5 journal placement, particularly when supported by causal inference methods, while mid-tier journals may exhibit less consistent priorities for novelty. Similarly, causal gap filling is more likely to be rewarded across publication tiers, whereas non-causal gap filling is less influential outside of top 5 journals. These findings underscore the importance of methodological rigor in shaping the reception of novel or underexplored research contributions.32

(Figure 13)

Conceptual Importance and Diversity

Figure 14 reports results for measures capturing conceptual importance and diversity, including average eigenvector or PageRank centralities, their variances, and the source–sink ratio of the knowledge graph.33 These measures provide insight into whether focusing on central or peripheral concepts, or maintaining a balanced versus skewed causal structure, correlates with publication success and citation impact.

Relationship between Conceptual Importance and Diversity Measures with Publication and Citations

Figure 14. Relationship between Conceptual Importance and Diversity Measures with Publication and Citations

Note: This figure shows coefficients for measures of conceptual importance (e.g., eigenvector centrality, PageRank), their variances, and the source-sink ratio. Results are presented for both “Non-Causal” and “Causal” versions of these measures, and for models with and without year fixed effects. Positive coefficients indicate associations with higher likelihood of publication in prestigious journals or with greater citation impact.

Papers that engage with more central concepts, as measured by higher mean PageRank and mean eigenvector centrality, tend to fare better in attracting citations. Both causal and non-causal versions of these measures are positively correlated with citation outcomes, though non-causal versions exhibit slightly larger coefficients. This suggests that engaging with well-established or influential concepts, even outside of a causal framework, enhances a paper’s longer-term academic impact.

The relationship between central concepts and publication outcomes is more mixed. Top 5 journals do not necessarily prefer highly central concepts; the coefficient for causal mean centrality is negative but statistically insignificant, while the coefficient for non-causal mean centrality is generally zero-to-positive and similarly statistically insignificant. This indicates that top-tier journals may prioritize novel or frontier topics over central, well-established concepts. Once published, however, papers centering on mainstream, well-known concepts tend to gain a stronger foothold in academic discourse, as reflected in their higher citation counts.

For top 6–20 and top 21–100 journals, non-causal mean centrality measures show positive correlations with publication outcomes, whereas causal versions exhibit coefficients close to zero. This highlights a potential preference for broader conceptual engagement without requiring rigorous causal identification at these journal tiers.

Variance in centralities provides insight into conceptual diversity. Greater diversity in conceptual centrality does not consistently raise the probability of top-tier publication but can sometimes benefit citations. Papers that mix central and peripheral concepts may appeal to diverse academic sub-audiences, enhancing long-term influence. However, for top 5 publications, higher variance in causal centralities tends to have a negative relationship, suggesting that highly diverse causal arguments may be viewed as less cohesive or rigorous by elite journals.

The source–sink ratio offers another perspective. For top 5 publications, a higher causal source–sink ratio (indicating more distinct causal sources converging on fewer outcomes) is positively correlated with publication outcomes, suggesting that balanced yet complex causal narratives are favored. In contrast, the coefficient for the non-causal source–sink ratio is negative, indicating that papers relying on non-causal sources for multiple effects are less likely to appear in top 5 journals. For citations, a higher causal source–sink ratio also exhibits a positive relationship, while the coefficient for non-causal source–sink ratio remains close to zero. For top 6–20 and top 21–100 publications, both causal and non-causal source–sink ratios exhibit coefficients close to zero, indicating limited relevance for these outcomes.

(Figure 14)

Conclusion

This paper introduces a graph-based approach to quantifying the structure and content of economic research, using a custom large language model to map both general and empirically supported claims in over 44,000 NBER and CEPR working papers. By analyzing the relationships between economic concepts and the methods used to substantiate claims, we establish new insights into how the discipline has evolved and how structural features of research predict academic reception.

First, the “credibility revolution” in economics has reshaped both how researchers establish claims and which claims gain traction. The proportion of empirically supported claims has risen dramatically, from about 4% in 1990 to nearly 28% in 2020, reflecting the growing emphasis on robust empirical identification methods. This shift reflects the increasing importance of credibility in shaping the direction and reception of economic research.

Second, the drivers of top-tier journal publication often differ from those of long-term citation impact. Journals such as the AER and QJE strongly reward the use of empirical methods, particularly when coupled with novel or underexplored concepts. Yet, once published, papers that engage with central and widely recognized topics tend to accrue more citations, revealing a persistent tension between pursuing frontier exploration to meet editorial priorities and addressing established domains to gain broader scholarly influence.

Third, our findings highlight the importance of narrative depth. We measure this using metrics such as the longest chain of reasoning in a paper, which are predictive of both top-tier publication and citations. In contrast, general narrative complexity unsupported by empirical claims often has no effect or is negatively correlated with these outcomes, suggesting that complexity alone is insufficient; it must be anchored in well-documented relationships.

Finally, bridging underexplored conceptual domains (“gap filling”) emerges as an important but a subtle factor. While editors reward efforts to connect previously neglected topics, these connections do not consistently lead to higher long-term citations unless they are situated within established debates or empirically supported frameworks.

Overall, our results highlight the evolving nature of academic evaluation in economics. Journals increasingly prioritize rigorous empirical support and innovative connections, yet broader citation impact often rewards alignment with recognized debates. Striking a balance between methodological rigor, novel contributions, and integration with central concepts appears essential for scholars seeking both visibility and lasting impact. This study offers a framework for understanding these trade-offs and navigating an increasingly competitive field, providing actionable insights into how economic knowledge develops and gains influence.

Figures and Tables

Source–Sink Ratio Over Time and by Field

Figure 9 (PDF p. 56). Source–Sink Ratio Over Time and by Field. (a) Time Trends (All vs. Causal). (b) Cross-Field Comparisons. Note: Panel (a) plots mean source–sink ratios over time (1980–2023) for both the non-causal subgraph (orange) and the causal subgraph (blue), with shaded areas representing 95% confidence intervals. The Causal series rises steadily, indicating that recent papers tend to incorporate more distinct causal drivers than outcomes. Panel (b) shows field-level average source–sink ratios, again splitting Non-Causal vs. Causal. Fields such as Behavioral and Health rank highly in the Causal version of the ratio, suggesting an emphasis on multiple identified factors leading to a smaller set of outcomes, whereas Econ History and Econometrics display relatively higher Non-Causal version of the ratios.

Appendix

Details on LLM-based Information Retrieval

Background on Large Language Models Large Language Models (LLMs) like GPT-4o-mini have significantly advanced natural language processing by enabling machines to understand and generate human-like text. Pre-trained on extensive datasets, including academic papers, books, websites, and other textual sources, these models capture the complexities of language, semantics, and context. This extensive pre-training allows LLMs to perform a variety of tasks, such as text summarization, translation, question answering, and information retrieval ([10], [30], [31], [46]).

In our study, we leverage the LLM’s ability to comprehend and extract structured information from unstructured text. By processing the first 30 pages of each economics working paper, the LLM identifies and extracts key metadata, methodological details, and causal claims. Unlike traditional NLP methods that rely on keyword matching or rule-based extraction, LLMs can understand complex language (often found in academic papers) and infer relationships between concepts, making them particularly effective for analyzing complex academic texts.

Prompt Design and Multi-Stage Extraction Process

Our information retrieval process involves a multi-stage approach, utilizing the LLM to extract and structure the necessary data efficiently.

Prompt Design

To effectively guide the LLM in extracting the required information, we developed comprehensive prompts that included detailed system and user instructions for each stage of the extraction process. The prompts were designed to specify the assistant’s role, define the task, provide definitions where relevant, enforce a structured response format, and set guidelines for accuracy and consistency.

Assistant’s Role and Task Definition We explicitly defined the assistant’s role as an expert in economics paper analysis, specializing in interpreting complex academic content and extracting nuanced information. This role specification was intended to align the LLM’s outputs with the expectations of an expert-level analysis. We provided detailed task definitions, instructing the assistant to analyze the text content of a provided economics paper and extract specific information related to metadata, causal claims, research methods, variables, data measurements, identification strategies, and data usage.

Inclusion of Canonical Definitions To ensure accurate and consistent classification of research methods, we included canonical definitions of key methodologies commonly used in economics, such as Randomized Controlled Trials (RCTs), Difference-in-Differences (DiD), Instrumental Variables (IV), Regression Discontinuity Designs (RDD), and others. By providing these standard definitions, we aimed to minimize ambiguity and enhance the reliability of method classification.

Structured Response Format We specified that the assistant’s response should adhere to a structured JSON schema with predefined fields and data types. This structured format was crucial for facilitating subsequent data processing, aggregation, and analysis across the large corpus of papers. The JSON schema included detailed definitions and instructions for each field to ensure accurate and consistent extraction. One of the instances where it works the best is to provide an array of causal edges in a paper. Other useful application is specifying the data type: boolean (e.g. whether paper uses any data), categorical (e.g., source of data: private or public sector), numeric (e.g., number of authors) or string response (e.g., name of the data provider).

Guidelines for Accuracy and Consistency We emphasized strict adherence to the schema, accuracy, and clarity in the assistant’s responses. We instructed the assistant to use the canonical definitions when classifying methods and to ensure that all required fields were accurately and completely filled out. These guidelines were essential to maintain the quality and consistency of the extracted data.

Overview of Prompts and Instructions

The following provides an overview of the system instructions provided to the LLM for each stage of the extraction process. Due to space constraints, we present summarized versions of the prompts. The complete system prompts and JSON schemas will be included in the replication package accompanying this paper.

Stage 1: Curated Summary In Stage 1, the assistant was instructed to analyze the first 30 pages of the paper and extract a curated summary of key elements. The system instructions included:

  • Assistant’s Role: An expert assistant specializing in analyzing economics papers.

  • Task: Extract specific information related to research questions, causal identification strategies, data usage, data accessibility, acknowledgements, and metadata.

  • Guidelines: Provide clear, detailed, and information-rich responses for each field. Focus exclusively on specified sections. Adhere strictly to specified formats and instructions.

  • Fields to Extract: Research questions from the abstract, introduction, and full text; causal identification information; causal claims; framing and policy implications; data and units of analysis; data accessibility; institutional and author-level information; acknowledgements.

Stage 2: Causal Graph Retrieval In Stage 2, the assistant was tasked with extracting detailed causal relationships from the summaries provided in Stage 1. The system instructions included:

  • Assistant’s Role: An expert assistant specializing in analyzing economics papers.

  • Task: Extract an exhaustive list of causal relationships from provided texts, capturing all intermediate steps, mediators, confounders, and other relevant nodes in the causal graph (DAG).

  • Guidelines: Be exhaustive in extraction. Use exact terms from the authors. Provide each causal relationship in a structured format with specified fields.

  • Fields to Extract: For each causal relationship, include the causal claim, cause, effect, type of causal relationship, whether evidence is provided, sign of impact, effect size, statistical significance, causal inference method, sources of exogenous variation, level of tentativeness.

Stage 3: Data and Accessibility In Stage 3, the assistant was instructed to extract structured information regarding data sources and accessibility from the data-related summaries in Stage 1. The system instructions included:

  • Assistant’s Role: An expert assistant specializing in analyzing economics papers.

  • Task: Extract specific pieces of information related to data and units of analysis, and data accessibility.

  • Guidelines: Carefully extract information for each field, adhering strictly to definitions and instructions. Use exact wording from the text when possible.

  • Fields to Extract: Data usage indicators, total number of observations, units of analysis, data granularity, temporal and geographical context, data ownership, data accessibility, ethical considerations.

Matching LLM Output to JEL Codes and OpenAlex Topics

After the LLM extracted the causal claims and provided free-text descriptions of the source and sink variables, we needed to standardize these variables to facilitate aggregation and systematic analysis across the corpus. This standardization was achieved by mapping the variable descriptions to official Journal of Economic Literature (JEL) codes and OpenAlex Topics.

Choice of Standardization Methods

We considered several options for standardizing the terms used in the source and sink variables, including JEL codes, OpenAlex Topics, Concepts, and the OECD Glossary of Statistical Terms. Each option had its advantages and limitations. JEL codes are well-understood within the economics community and facilitate interpretability but are relatively broad in classification. OpenAlex Topics offer more granularity with around 4,500 topics but are less familiar to economists. Concepts provide even more detail with approximately 60,000 terms but are being deprecated by OpenAlex and are not widely recognized in economics. The OECD Glossary contains about 6,700 economics-related terms but may be biased towards statistical concepts.

After careful consideration, we opted to focus on JEL codes for standardization. JEL codes were chosen as our primary method due to their familiarity and acceptance within the economics discipline, facilitating interpretation and communication of results.34

Embedding-Based Matching Methodology

To map the free-text variable descriptions to the standardized codes, we employed an embedding-based matching approach using vector representations of the texts. We utilized the OpenAI text embedding model text-embedding-3 large, which generates 1,024-dimensional vector embeddings that capture the semantic meaning of the text. Embeddings were generated for: (i) the free-text descriptions of the source and sink variables extracted by the LLM, (ii) the official descriptions of JEL codes, enhanced by concatenating descriptions from higher-level codes to provide richer context, and (iii) the summaries of OpenAlex Topics.

We calculated the cosine similarity between the embeddings of the variable descriptions and the embeddings of the JEL codes and OpenAlex Topics.35 We assigned the variable to the codes/topics with the highest similarity scores.

Advantages and Limitations of the Embedding-Based Approach

By leveraging embeddings, we moved beyond simple keyword matching, which can be limited by variations in terminology and phrasing. The embedding-based approach captures the semantic meaning of the texts, allowing us to match variable descriptions to standardized concepts even when different terms are used to describe similar ideas (e.g., unemployment rate” vs. joblessness”). This method enhanced the robustness of our matching process, reducing the impact of typos, synonyms, and variations in language. It allowed us to systematically standardize a large number of variable descriptions across the corpus, facilitating cross-paper comparisons and aggregations.

While the embedding-based matching approach offers significant advantages, there are limitations to consider. The quality of the matches depends on the accuracy of the embeddings and the chosen similarity threshold, which in our case is the one with highest similarity. By focusing on only one matching concept, we are imposing a structure on the latent causal graph. It may well be that a source or a sink could be captured by multiple JEL codes. However, for simplicity and consistency across papers, we decided to capture only the best match. Future research can consider setting a threshold, e.g., 0.85, beyond which multiple JEL codes can be accepted. However, this comes with its own set of hyper-parameter selection: there is a trade-off between precision and recall; a higher threshold increases precision but may miss relevant matches, while a lower threshold increases recall but may include irrelevant matches.36

By using both JEL codes and OpenAlex Topics, we leveraged the strengths of each system, with JEL codes providing familiarity and OpenAlex Topics offering granularity. Overall, the use of embeddings was instrumental in standardizing and operationalizing our source and sink variables.

Validation and Quality Assurance

To ensure the reliability of the extracted data, we implemented validation checks at each stage. The structured outputs were parsed and checked for compliance with the specified schemas. In cases where inconsistencies or missing data were detected, the prompts were refined, and the extraction was repeated.

We also conducted manual reviews of a sample of the extracted data to assess accuracy. This included checking the correctness of the causal claims extracted, the appropriateness of the mapped JEL codes, and the consistency of metadata. The feedback from these reviews informed further refinements to the prompts and extraction process.

A large scale validation of our approach will follow in future drafts. This includes contacting corresponding authors to validate the causal graphs in their own papers.

Limitations and Considerations

While the use of LLMs provides significant advantages in processing large volumes of complex text, there are limitations to consider. The LLM’s extraction is dependent on the quality and clarity of the original text; ambiguities or omissions in the papers may lead to incomplete extraction. The LLM may occasionally misclassify or misinterpret information, particularly with nuanced methodological details. While we collected additional attributes such as effect sizes and statistical significance, these features were experimental and not used in the main analysis due to variation in reporting standards.

Additionally, the embedding-based matching approach for standardizing variables may not capture all nuances of the economic concepts involved. There is a risk of misclassification if the variable descriptions are ambiguous or if the embeddings do not accurately represent the semantic content. Despite these limitations, we believe that the overall methodology provides a robust framework for large-scale analysis of economic research.

Replication Package

The full system prompts, JSON schemas, and code used in the extraction process will be included in the replication package accompanying this paper. Researchers interested in replicating or extending our analysis can refer to these materials for detailed guidance.37

Validating information retrieval

Validation with the [18] Dataset

To validate the accuracy of our AI-based information retrieval methods, we compared our dataset with an external dataset from [18], which examines p-hacking, data type, and data-sharing policies in economics. Their dataset includes 1,106 articles published in leading economics journals between 2002 and 2020, with detailed annotations on the empirical methods used and the fields of study.

We matched our dataset with theirs using paper titles, employing both direct matches and fuzzy matching techniques to maximize the number of matched papers. In total, we matched 307 papers between the two datasets. For these matched papers, we compared our classifications of empirical methods—Difference-in-Differences (DiD), Randomized Controlled Trials (RCT), Regression Discontinuity Design (RDD), Instrumental Variables (IV)—and fields—Urban Economics, Finance, Macroeconomics, Development—with those from [18] to assess the accuracy of our information retrieval.

We calculated two evaluation metrics: Accuracy and the F1 Score, where the ground truth labels are those from [18]. The results are presented in Table A2.

VariableAccuracyF1 Score
Method: DiD0.77620.8447
Method: RCT0.70630.8269
Method: RDD0.93710.9644
Method: IV0.74130.8183
Field: Urban Economics0.91380.9545
Field: Finance0.87880.9295
Field: Macroeconomics0.97440.9870
Field: Development0.66430.7937

Table A2. Validation Results of Information Retrieval

The results indicate that our information retrieval methods achieve high levels of accuracy and F1 scores in identifying both empirical methods and fields of study. In particular, the identification of RDD methods and classification of papers in Macroeconomics and Urban Economics show excellent performance, with accuracy and F1 scores exceeding 0.90. These findings provide confidence in the reliability of our AI-based extraction and classification processes in at least two dimensions of our analysis: fields and methods.

Validation with the Plausibly Exogenous Galore Dataset

To further validate our information retrieval methods, we compared our data on causal claims with an external dataset, the Plausibly Exogenous Galore dataset38. This dataset includes entries for 1,435 papers (as of August 2024), each with the main Left-Hand Side (LHS) and Right-Hand Side (RHS) variables, as well as the primary source of exogenous variation, as identified by authors and contributors. Unlike our dataset, which captures causal claims at the paper-claim level, the Plausibly Exogenous Galore dataset records information at the paper level, often focusing on key variables rather than the complete knowledge graph of causes, effects, and sources.

We matched 485 papers from our dataset to entries in the Plausibly Exogenous Galore dataset using exact and fuzzy title matching. To facilitate a meaningful comparison, we aggregated our data at the paper level by concatenating all causes, effects, and sources of exogenous variation. This resulted in structured phrases: “The causes are: <cause>; <cause>; ...” for causes, “The effects are: <effect>; <effect>; ...” for effects, and “The source(s) of exogenous variation are: <source>; <source>; ...” for sources. Similarly, the Plausibly Exogenous Galore data was formatted as “The cause is: <RHS>” for the RHS (cause), “The effect is: <LHS>” for the LHS (effect), and “The source(s) of exogenous variation are: <source of exogenous variation>.”

Using OpenAI’s text-embedding-3-large embeddings model, we generated embeddings for each component separately (causes, effects, and sources of exogenous variation) and calculated cosine similarity scores for each component between our data and the Plausibly Exogenous Galore dataset. We computed the average, minimum, maximum, median, and standard deviation of these similarities.

The results (Table A3) show that while cause and effect similarities were moderate, reflecting variability in how these are represented across claims within a paper, the source of exogenous variation achieved a higher similarity score. This result supports the reliability of our method for capturing consistent exogenous variation sources, as these tend to vary less across claims. These findings reinforce our information retrieval system’s ability to align with external validations, particularly for core causal identifiers.

VariableMean SimilarityMedianStd DevMinMax
Cause Similarity0.61400.62450.11270.22850.8798
Effect Similarity0.63860.64670.08930.37950.8452
Exogenous Variation Similarity0.80140.81420.09750.49200.9863

Table A3. Validation Results with Plausibly Exogenous Galore Dataset

Appendix Tables and Figures

Proportion of Papers Published in Top 5 Journals, Pre- and Post-2000

Figure A1. Proportion of Papers Published in Top 5 Journals, Pre- and Post-2000. (a) By Field. (b) By Field and Method. Note: This figure displays the proportion of working papers that were eventually published in the top five economics journals, broken down by (a) field and (b) field by empirical method. Data are derived from our matched publication dataset. Arrows in panel (a) indicate changes between pre-2000 and post-2000 periods. The heatmap in panel (b) shows that certain fields and methods have higher publication rates in top journals, with field-method combinations such as Theoretical methods in Behavioural, Structural in IO or RCTs in Urban.

Distribution of citation percentiles by journal category

Figure A2. Distribution of Citation Percentiles by Journal Category

Note: This figure displays kernel density plots of the citation percentiles for papers published in Top 5, Top 6–20, and Top 21–100 journals. The citation percentiles are calculated based on the entire sample, with higher values indicating higher citation counts relative to other papers. The plot shows that while papers published in higher-ranked journals tend to receive more citations on average, the most highly cited papers are more evenly distributed across journal categories. This suggests that exceptionally influential papers can emerge from a wide range of journals.

Journal CategoryMean PercentileMedian PercentileStandard Deviation
Top 562.186825.86
Top 6–2057.066127.97
Top 21–10052.125327.37

Table A1. Summary Statistics of Citation Percentiles by Journal Category

Note: This table summarizes the citation percentiles for papers published in different journal categories. The mean and median percentiles indicate that papers published in Top 5 journals have higher citation impact on average compared to those in lower-ranked journals. However, the overlap in distributions suggests that highly cited papers can also be found in lower-ranked journals.

Note: This table presents the validation results of our information retrieval methods compared to the ground truth provided by [18]. Variable indicates either the empirical method (DiD, RCT, RDD, IV) or the field of study (Urban Economics, Finance, Macroeconomics, Development). Accuracy measures the proportion of correctly classified instances, while F1 Score represents the harmonic mean of precision and recall. Higher values indicate better performance.

Note: This table presents the cosine similarity results between our dataset and the Plausibly Exogenous Galore dataset, which records plausibly exogenous sources, main causal factors, and outcomes in the economics literature. Mean Similarity indicates the average cosine similarity score, while Std Dev shows the standard deviation of these scores, capturing the variability. Min and Max represent the range of similarity scores across matched papers.

References

  1. [1]Abramitzky, R., Greska, L., Pérez, S., Price, J., Schwarz, C. and Waldinger, F. (2024), Climbing the ivory tower: How socio-economic background shapes academia, Technical report, National Bureau of Economic Research.
  2. [2]Airoldi, A. and Moser, P. (2024), Inequality in science: Who becomes a star?, Technical report, National Bureau of Economic Research.
  3. [3]Alabrese, E. (2022), ‘Bad science: Retractions and media coverage’.DOI
  4. [4]Alabrese, E., Capozza, F. and Garg, P. (2024), ‘Politicized scientists: Credibility cost of political expression on twitter’.
  5. [5]Andre, P. and Falk, A. (2021), What’s worth knowing in economics? a global survey among economists. Working Paper.
  6. [6]Angrist, J., Azoulay, P., Ellison, G., Hill, R. and Lu, S. F. (2017), ‘Economic research evolves: Fields and styles’, American Economic Review: Papers and Proceedings 107(5), 293–297.DOI
  7. [7]Angrist, J. D. and Pischke, J.-S. (2008), Mostly Harmless Econometrics: An Empiricist’s Companion, Princeton University Press, Princeton, NJ.
  8. [8]Angrist, J. D. and Pischke, J.-S. (2010), ‘The credibility revolution in empirical economics: How better research design is taking the con out of econometrics’, Journal of Economic Perspectives 24(2), 3–30.
  9. [9]Angrist, J. and Imbens, G. (1994), ‘Identification and estimation of local average treatment effects’.
  10. [10]Ash, E. and Hansen, S. (2023), ‘Text algorithms in economics’, Annu. Rev. Econom. .DOI
  11. [11]Athey, S. and Imbens, G. W. (2019), ‘Machine learning methods that economists should know about’, Annual Review of Economics 11(1), 685–725.DOI
  12. [13]Banerjee, A., Duflo, E., Glennerster, R. and Kinnan, C. (2015), ‘The miracle of microfinance? evidence from a randomized evaluation’, American economic journal: Applied economics 7(1), 22–53.DOI
  13. [14]Baumann, A. and Wohlrabe, K. (2020), ‘Where have all the working papers gone? evidence from four major economics working paper series’, Scientometrics 124(3), 2433–2441.
  14. [15]Behavioural Insights Team (2024), ‘A blueprint for better international collaboration on evidence’. URL: provide URL here
  15. [16]Bloom, N., Jones, C. I., Van Reenen, J. and Webb, M. (2020), ‘Are ideas getting harder to find?’, American Economic Review 110(4), 1104–1144.
  16. [17]Bonacich, P. (1972), ‘Technique for analyzing overlapping memberships’, Sociological methodology 4, 176–185.DOI
  17. [18]Brodeur, A., Cook, N. and Neisser, C. (2024), ‘P-hacking, data type and data-sharing policy’, The Economic Journal 134(659), 985–1018.
  18. [19]Buehler, M. J. (2024), ‘Accelerating Scientific Discovery with Generative Knowledge Extraction, Graph-Based Representation, and Multimodal Intelligent Graph Reasoning’, Machine Learning: Science and Technology 5(4), 045005. URL: https://iopscience.iop.org/article/10.1088/2632-2153/ad7228arxiv.org/abs/2403.11996
  19. [20]Card, D. and DellaVigna, S. (2013), ‘Nine facts about top journals in economics’, Journal of Economic Literature 51(1), 144–161.
  20. [21]Cartwright, N. (2007), ‘Are rcts the gold standard?’, BioSocieties 2(1), 11–20.DOI
  21. [22]Chan, J., Akamatsu, M., Vargas, D., Kawerau, L. and Gartner, M. (2024), ‘Steps towards an infrastructure for scholarly synthesis’, arXiv preprint arXiv:2407.20666.arxiv.org/abs/2407.20666
  22. [23]Chernozhukov, V., Chetverikov, D., Demirer, M., Duflo, E., Hansen, C., Newey, W. and Robins, J. M. (2018), ‘Double/debiased machine learning for treatment and structural parameters’, The Econometrics Journal 21(1), C1–C68.
  23. [24]Chetty, R., Hendren, N., Kline, P. and Saez, E. (2014), ‘Where is the land of opportunity? the geography of intergenerational mobility in the united states’, The quarterly journal of economics 129(4), 1553–1623.
  24. [25]Currie, J., Kleven, H. and Zwiers, E. (2020), Technology and big data are changing economics: Mining text to track methods, in ‘AEA Papers and Proceedings’, Vol. 110, American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203, pp. 42–48.
  25. [26]Davies, B. (2022), ‘Gender sorting among economists: Evidence from the nber’, Economics Letters 217, 110640.DOI
  26. [27]de Quidt, J., Haushofer, J. and Roth, C. (2018), ‘Measuring and bounding experimenter demand’, American Economic Review 108(11), 3266–3302. URL: https://www.aeaweb.org/articles?id=10.1257/aer.20171330
  27. [28]Deaton, A. (2010), ‘Instruments, randomization, and learning about development’, Journal of Economic Literature 48(2), 424–455.DOI
  28. [29]Deaton, A. and Cartwright, N. (2018), ‘Understanding and misunderstanding randomized controlled trials’, Social Science & Medicine 210, 2–21.
  29. [30]Dell, M. (2024), ‘Deep learning for economists’, National Bureau of Economic Research Working Paper Series (32768). URL: http://www.nber.org/papers/w32768
  30. [31]Fetzer, T., Lambert, J. P., Garg, P. and Feld, B. (2024), Ai-generated production networks: Measurement and applications to global trade, Technical report, Working Paper.DOI
  31. [32]Gabaix, X. (2011), ‘The granular origins of aggregate fluctuations’, Econometrica 79(3), 733–772.
  32. [33]Garg, P. and Fetzer, T. (2024), ‘Political expression of academics on social media’.DOI
  33. [34]Gelman, A. and Loken, E. (2014), ‘The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition” or “p-hacking” and the research hypothesis was posited ahead of time’, Department of Statistics, Columbia University 348, 1–17.
  34. [35]Goldberg, P. K., Khandelwal, A. K., Pavcnik, N. and Topalova, P. (2010), ‘Imported intermediate inputs and domestic product growth: Evidence from india’, The Quarterly journal of economics 125(4), 1727–1767.
  35. [36]Goldsmith-Pinkham, P. (2024), ‘Tracking the credibility revolution across fields’, arXiv preprint arXiv:2405.20604 .
  36. [37]Hamermesh, D. S. (2013), ‘Six decades of top economics publishing: Who and how?’, Journal of Economic Literature 51(1), 162–172.
  37. [38]Heckman, J. J. (2001), ‘Micro data, heterogeneity, and the evaluation of public policy: Nobel lecture’, Journal of Political Economy 109(4), 673–748.
  38. [39]Hidalgo, C. A. and Hausmann, R. (2009), ‘The building blocks of economic complexity’, Proceedings of the National Academy of Sciences 106(26), 10570–10575. URL: https://www.pnas.org/doi/abs/10.1073/pnas.0900943106
  39. [40]Imbens, G. W. and Rubin, D. B. (2015), Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction, Cambridge University Press.
  40. [41]Ioannidis, J. P. (2005), ‘Why most published research findings are false’, PLOS Medicine 2(8), e124.
  41. [42]Ioannidis, J. P. (2016), ‘The mass production of redundant, misleading, and conflicted systematic reviews and meta-analyses’, The Milbank Quarterly 94(3), 485–514.DOI
  42. [43]Ioannidis, J. P., Stanley, T. D. and Doucouliagos, H. (2017), ‘The power of bias in economics research’.DOI
  43. [44]Keane, M. P. (2010), ‘A structural perspective on the experimentalist school’, Journal of Economic Perspectives 24(2), 47–58.DOI
  44. [45]Koffi, M., Pongou, R. and Wantchekon, L. (2024), The color of ideas: Racial dynamics and citations in economics, Technical report, National Bureau of Economic Research.
  45. [46]Korinek, A. (2023), ‘Generative ai for economic research: Use cases and implications for economists’, Journal of Economic Literature 61(4), 1281–1317.DOI
  46. [47]Leskovec, J., Lang, K. J., Dasgupta, A. and Mahoney, M. W. (2008), Statistical properties of community structure in large social and information networks, in ‘Proceedings of the 17th international conference on World Wide Web’, pp. 695–704.DOI
  47. [48]Mellon, J. (2021), ‘Rain, rain, go away: 194 potential exclusion-restriction violations for studies using weather as an instrumental variable’, American Journal of Political Science .
  48. [49]Merton, R. K. (1968), ‘The matthew effect in science’, Science 159(3810), 56–63.DOI
  49. [50]Milo, R., Shen-Orr, S., Itzkovitz, S., Kashtan, N., Chklovskii, D. and Alon, U. (2002), ‘Network motifs: simple building blocks of complex networks’, Science 298(5594), 824–827.
  50. [51]Nature Editorial (2024), ‘Unearthing ‘hidden’ science would help to tackle the world’s biggest problems’, Nature 633(7930), 493. Editorial. URL: https://doi.org/10.1038/d41586-024-02991-5
  51. [52]Page, L. (1999), The pagerank citation ranking: Bringing order to the web, Technical report, Technical Report.
  52. [53]Park, M., Leahey, E. and Funk, R. J. (2023), ‘Papers and patents are becoming less disruptive over time’, Nature 613(7942), 138–144.DOI
  53. [54]Pearson, H. (2024), ‘Can ai review the scientific literature—and figure out what it all means?’, Nature 635(8038), 276–278.DOI
  54. [55]Ravallion, M. (2009), ‘Should the randomized controlled trial be the gold standard for development research?’, World Bank Research Observer 24(1), 30–53.
  55. [56]Rawat, S. and Meena, S. (2014), ‘Publish or perish: Where are we heading?’, Journal of research in medical sciences: the official journal of Isfahan University of Medical Sciences 19(2), 87.pubmed.ncbi.nlm.nih.gov/24778659
  56. [57]Resnik, D. B. (1998), The Ethics of Science: An Introduction, Routledge, London, UK.
  57. [58]Rosenbaum, P. R. and Rosenbaum, P. R. (2002), Overt bias in observational studies, Springer.
  58. [59]Rubin, D. B. (1984), ‘Bayesianly justifiable and relevant frequency calculations for the applied statistician’, Annals of Statistics 12(4), 1151–1172.DOI
  59. [60]Shen-Orr, S. S., Milo, R., Mangan, S. and Alon, U. (2002), ‘Network motifs in the transcriptional regulation network of escherichia coli’, Nature genetics 31(1), 64–68.
  60. [61]Simonsohn, U., Nelson, L. D. and Simmons, J. P. (2014), ‘p-curve and effect size: Correcting for publication bias using only significant results’, Perspectives on Psychological Science 9(6), 666–681.
  61. [62]Sims, C. A. (2010), ‘But economics is not an experimental science’, Journal of Economic Perspectives 24(2), 59–68.DOI
  62. [63]Small, H. (1973), ‘Co-citation in the scientific literature: A new measure of the relationship between two documents’, Journal of the American Society for information Science 24(4), 265–269.DOI
  63. [64]Van Eck, N. J. and Waltman, L. (2014), Visualizing bibliometric networks, in ‘Measuring scholarly impact: Methods and practice’, Springer, pp. 285–320.DOI
  64. [65]Waltman, L. and Van Eck, N. J. (2012), ‘A new methodology for constructing a publication-level classification system of science’, Journal of the American Society for Information Science and Technology 63(12), 2378–2392.arxiv.org/abs/1203.0532
  65. [66]Wasserstein, R. L. and Lazar, N. A. (2016), ‘The asa’s statement on p-values: Context, process, and purpose’, The American Statistician 70(2), 129–133.DOI
  66. [67]Weitzman, M. L. (1998), ‘Recombinant growth’, The Quarterly Journal of Economics 113(2), 331–360.DOI

Paper details

Contents