Education.org

Published on January 22, 2023

Why is so much useful education evidence still overlooked?

Lessons from global initiatives on tapping into grey literature and practitioner insights

Image for Why is so much useful education evidence still overlooked?

Efforts to improve education systems globally increasingly depend on evidence-informed policymaking. Yet much of the knowledge that could strengthen decision-making—particularly in low- and middle-income countries—remains underutilised. This includes grey literature, practitioner knowledge, and qualitative insights that are often excluded by traditional definitions of “good evidence.”

This report challenges prevailing hierarchies of evidence and explores how to expand what counts as credible and useful for education policy. Drawing on a global review of 26 initiatives, it highlights how governments, research organisations, and international partners are working to identify, appraise, and integrate diverse types of evidence into policy processes.

Key findings include:

  • Expanding inclusion: Initiatives are making greater use of grey literature, experiential knowledge, and context-rich qualitative data, especially where formal academic research is limited or slow to emerge.
  • Challenging hierarchies: Some actors are beginning to rethink rigid evidence hierarchies that privilege randomised trials and peer-reviewed studies, instead adopting more pragmatic or equity-focused approaches to evidence use.
  • Developing new tools: A range of methods and frameworks are emerging to help assess the quality, relevance, and usability of non-traditional evidence sources.
  • Making sense of abundance: With large volumes of literature available, many initiatives are developing typologies and classification systems to help users navigate and extract insights efficiently.
  • Bridging supply and demand: Broader definitions of evidence are helping to close the gap between what researchers produce and what policymakers need—often more contextualised, timely, and actionable information.

The report argues that broadening the scope of evidence is not just a technical task but a matter of inclusion and legitimacy. By valuing a wider range of knowledge, education systems can become more responsive to the lived realities of teachers, learners, and communities—ultimately leading to better, more equitable outcomes.

It concludes with practical guidance for those designing evidence infrastructures or supporting policymaking in low-resource contexts, offering principles and illustrative tools for more inclusive evidence use.

Why is this review needed? Understanding how to make education evidence more inclusive

1.1 Rationale

Although investments in education research and initiatives have increased in recent decades, there has been limited progress in how this evidence is used in real-world decision-making. The evidence available to education leaders remains fragmented, hard to access, and often lacks relevance to specific contexts.

Frontline education leaders are asking for better support to navigate the large and growing volume of education evidence published each year. They also want access to practical, locally grounded knowledge that often goes untapped. While decisionmakers seek evidence that is timely and actionable, much of the existing research is geared toward academic standards rather than the practical needs of educators and planners.

A major portion of this overlooked material—often called grey literature—includes insights from programme implementation, policy experience, and practitioner knowledge. It’s not always formal research. It may be descriptive reports, internal reviews, or firsthand accounts. Increasing the use of evidence in education policy, planning, and practice means expanding the types of evidence that are recognised and valued.

This report provides background analysis to support an international initiative—led by USAID and Education.org—focused on making education evidence more inclusive. This effort, known as the International Working Group on Inclusive Evidence in Education (IWG), brings together a diverse group of stakeholders—researchers, evidence synthesizers, policymakers, practitioners, and funders. The goal is to widen the range of evidence considered in decision-making, particularly by drawing from sources beyond academic journals and commercial publications.

While peer-reviewed research has established norms for quality and bias, such standards are often lacking in grey literature. One of IWG’s aims is to explore how these gaps can be addressed, and how more locally generated knowledge—often excluded from traditional systems—can be recognised and used more effectively.

The initiative supports Education.org’s broader mission to help education leaders by translating and synthesising a wide range of evidence, and by building bridges between those who produce knowledge and those who use it.

This review – or ‘landscaping exercise’ - also builds on Education.org’s earlier work. The organisation’s white paper, Calling for an Education Knowledge Bridge, set out a critique of the gap between research and practice in education. Work during the COVID-19 recovery period underscored the importance—and difficulty—of accessing context-specific, practitioner-informed evidence. In a recent prototype synthesis on Accelerated Education Programmes (AEP), Education.org tested new ways of gathering evidence, including crowdsourcing. This led to 76% of the final sources coming from unpublished or grey literature, which greatly enriched understanding of how AEPs work in practice. That review helped decisionmakers navigate both planning and implementation challenges more effectively.

This current review revisits and builds on that earlier work.

1.2 Purpose

The purpose of this review is to learn from existing initiatives that are already working to make evidence more inclusive—particularly by incorporating both published and unpublished sources. We examined how these initiatives approach four key steps:

  1. Defining what counts as evidence
  2. Finding and collecting relevant material
  3. Classifying different types of evidence
  4. Assessing the quality and relevance of that evidence

These are the same steps followed in many evidence synthesis or review processes. Our aim was to find practical ways to make it easier to identify and use grey literature, and to explore how simple appraisal principles can improve how this type of evidence is used in decision-making.

Because grey literature often includes qualitative information, we paid special attention to how initiatives include and assess qualitative evidence. While the challenge of collecting grey literature efficiently remains, several initiatives offer promising strategies to make the process more systematic and scalable over time.

This review also provides insights for how Education.org might improve its own approach to evidence synthesis in the future. These are summarised in the final section of the report and include ideas for presenting findings (e.g., using quality scales) and a curated list of databases for grey literature.

Methodology

2.1 Design stage

The driving question for the landscaping exercise was defined first. The exercise has the dual scope of examining initiatives for lessons in bringing evidence that is not published through traditional channels to the attention of decisionmakers as well as to identify lessons for future work of Education.org on producing evidence syntheses. Sub-questions for the analysis were broken into fields in a data collection form in Excel. These include defining sources, selection criterion, inclusion of unpublished works, tools or rubrics for appraising studies, and whether they monitor the use of unpublished sources.

Criterion for the selection of organisations and initiatives were set, drawing on the White Paper analysis, recommendations from partners, the team’s own familiarity, and web research. Initiatives (or organisations) were purposefully drawn from inside and outside the education sector. A consultant with strong health evidence experiences, but no education sector background, complemented the education experts within Education.org. Most initiatives were chosen because they are actively involved in evidence reviews or evidence syntheses, regardless of the sector. This approach was chosen because the universe of initiatives involved in original research is vast and it was not feasible to screen individual research initiatives to determine if they accessed grey literature or tried to integrate an inclusive set of evidence. A final step of the design stage was to develop a project charter. A briefing was held for all Education.org staff on the purpose, methods, and end products planned.

2.2 Data collection stages

Three overlapping data collection stages were designed: web research; document research; and follow-up queries or interviews. Five people were part of the data collection stages. For the web research, a trial of the data collection form was held to address inter-rater reliability issues. Each team member reviewed two assigned organisations or initiatives and filled out the data collection form. Results were discussed and the form and instructions were adjusted.

The team conducted web research for remaining priority organisations, filling out the form for each organisation or initiative. Weekly discussion sessions were held among team members to share what was being learned and on which fields to do more probing. As the figure indicates, 53 organisations and initiatives were on the most comprehensive list. Several organisations were found to be involved in original research only and were set aside as not meeting the criteria. In all, 27 were discarded for the reasons indicated in the figure. The websites of 26 were then studied in depth (see Annex 1 for this list). While health sector initiatives dominated (13), seven were found to focus on education alone, two on multiple sectors including education, and another four on other sectors.

Body image

Figure: Number of organisations and initiatives producing evidence reviews that were screened for guidance on inclusive evidence 

The web research identified multiple key reports, guidelines, or other documents. These were shared across the team for review to draw out lessons for applying to grey literature at three stages (1) identifying and collecting evidence; (2) classifying the evidence, and (3) appraising the quality of evidence for inclusion. Several lists of principles were compiled, and a summary of each document was created.

A separate desk study on Citizen Science was conducted to explore additional lessons for defining acceptable evidence and exploring lessons for participatory means of collecting evidence.

Initially, a more systematic literature review was envisioned and an effort to define search parameters was attempted during the data collection phase. However, the effort was rethought as confidence grew that we were finding most the relevant documents through the web search. Therefore, the document review phase focused on a purposeful sample that includes guidelines for evidence reviews and syntheses as well as guidelines for individual research studies, such Building Evidence in Education (BE2)1 Guidance Notes.

The web research and document review identified a few gaps or areas to probe through direct outreach in a small subset of initiatives examined. These queries, through email or calls, aim to enhance the analysis. Some of these calls also feed into the profiles of IWG members, or even specific candidates. A separate set of calls will continue to be conducted to build external understanding and support for the inclusive evidence in education initiative by updating these actors and soliciting ideas. Calls to date have been held with USAID’s learning network SHARE, INEE coordinator for the 2023 Data and Evidence Summit, Dr Wiggs at the Innovation Project, and Professor Reimers at the Harvard Graduate School of Education. Many more are being scheduled and as we reach out, people have been excited about this agenda and speak of its importance.

2.3 Analysis

The analytic writing phase drew from the Excel data file and reflection memos prepared by the consultant and the project lead. An additional pass was made of each of the documents reviewed to pull out more detail or extract quotations. Probing interviews were started in January and will influence the revised text for the IWG design document, rather than this landscaping report.

2.4 Limitations

The methodology followed for this landscaping exercise does not adequately capture the perspectives of knowledge producers and brokers operating at the national and regional levels. Outreach efforts to key local knowledge producers, including education research networks, in Kenya and Sierra Leone will continue. These consultations are expected to inform the June agenda and help us identify strong IWG members.

Findings and implications for advancing evidence inclusiveness

3.1 Dominant practices

This landscaping exercise confirmed several dominant characteristics of existing approaches to conducting evidence reviews and syntheses. While few of the approaches recognise unpublished evidence, the existing practices at multiple stages can be modified, through the efforts of the IWG, to facilitate and legitimise unpublished or informally published evidence that better serves education decisionmakers.

The discussion below examines what we did and did not find regarding audience, choice of topics, questions being addressed, inclusion of grey literature, diversity of voices, methods for assessing trustworthiness, and metric for monitoring progress in evidence accessibility or evidence inclusiveness. A number of these issues are discussed in greater detail in subsequent sections.

  • Audience: We found that many initiatives conducting systematic reviews or evidence syntheses claim to serve policy maker audiences in their systematic reviews. However, most do not address the kinds of questions senior government officials have and generally are not focused on the pragmatic issues of priority to policy makers. 
  • The choice of topic: This is generally made in response to academic or funder priorities. Several initiatives indicate that policy makers are involved in prioritising the topic.
  • The questions being asked: The review questions drive the protocols that define the sources to search, the type of evidence to be selected, etc. Most evidence reviews and syntheses do not attempt to answer questions of why something works in the contexts that it did, and what are the choices to be made in design and implementation. More common is a focus on the observed effect in controlled environments. As a result, the protocols2 for most evidence syntheses set methodologically rigid (for example only randomised control trials) definitions for the evidence to be used. They are not responsive to the context and exclude important information-rich respondents or methods that might yield more useful evidence.
  • Acceptance of grey literature: Nine of the 26 initiatives reviewed include grey literature research in the definition of evidence they include, primarily by searching databases such as OpenGrey, National Technical Information Service, and dissertation databases as well as including reports and conference abstracts. Several initiatives contact relevant individuals and organizations for information about unpublished or ongoing studies. We found no documentation of initiatives physically collecting grey literature directly from sources; generally sourcing is a desk-based process.
  • Inclusion of diverse voices: Rarely is there an explicit concern for the diversity of stakeholders represented in the data. When beneficiaries (policy makers, patients, practitioners) are mentioned as part of the evidence review, their roles are limited to involvement in prioritisation, design, implementation, and dissemination, not as experts themselves offering evidence. CUREE and the EPPI-Centre may also involve stakeholders in interpreting findings. The EPPI-Centre allows people’s views on their needs or requirements and descriptions of how or whether current policies are being implemented. The Clearing House for Military Family Readiness may include anecdotal data such as participants’ comments and satisfaction.
  • Assessing trustworthiness of evidence: Almost all initiatives are explicit in how they assess the trustworthiness of the evidence being considered by being attentive to minimising bias. Twenty out of the 26 initiatives have explicit procedures and standards, although some vary with the type of review or synthesis being undertaken. The Campbell Collaboration has a typical requirement for their inclusion: "Use (at least) two people working independently to apply a risk of bias/study quality tool or coding scheme to each included study and define in advance the process for resolving disagreements." Six out of 26 are unclear, at least not online or in publications, about their a priori criterion for finding and appraising evidence.
  • Metrics for monitoring progress: We found no lessons for monitoring progress on making evidence more accessible. Nor did we find efforts to monitor progress in integrating inclusive evidence. This is because there are so few initiatives trying to make evidence more inclusive. In addition, most organisations do not put their Key Performance Indicators online.

3.2 Alternative perspectives

Rejecting hierarchies

Almost twenty years ago, Petticrew and Roberts (2003) argued that typologies rather than a hierarchy of evidence are more useful in appraising evidence for social interventions, such as public health. They proposed a matrix approach to match questions to specific types of research. In education research circles, there is growing recognition that a hierarchy of designs and methods is not useful. The Building Evidence in Education working group (BE2 2015: 5) uses a similar argument that the context and research questions determine the appropriateness of a design and methods, as does the availability of data.

Questioning peer review

Many initiatives require that the evidence included is peer reviewed. It is interesting to note that a multi-faceted critique of the 60-year-old “sacred” peer review process is emerging, highlighting the poor job done by reviewers, the censoring role of the process, among other problematic aspects (Mastroianni 2022).

Defining best available evidence

Two particularly useful initiatives are highlighted below, the Joanna Briggs Institute as it addresses text and opinion evidence, and the Ways of Evaluating Important and Relevant Data’ (WEIRD) tool.

The Joanna Briggs Institute (JBI) stresses the importance of using “the best available evidence” and that may mean including text and opinion in systematic reviews, especially to inform clinical decision making (McArthur, et al., 2020).

Expert opinion has a role to play in evidence-based health care, as it can be used to either complement empirical evidence or, in the absence of research studies, stand alone as the best available evidence. While rightly claimed not to be a product of ‘good’ science, expert opinion is empirically derived and mediated through the cognitive processes of practitioners who have been typically trained in scientific method. This is not to say that the superior quality of evidence derived from rigorous research is to be denied; rather, that in its absence, it is not appropriate to discount expert opinion as non-evidence.

It is notable that “expert” is defined as practitioner, which can be applied in education, with recognition that not all education practitioners are trained in the scientific method.

JBI provides guidance on using text and opinion evidence to support “not only the evidence on the effectiveness of interventions (“knowing what” type of evidence), but also evidence related to subjective human experiences, culture, values, ethics, health policy, or the accepted discourse at the time of practice (“knowing how” type of evidence)” (citing Jordan, Konno & Mu, 2011). Sources of this non-research evidence include “expert opinions, consensus, current discourse, comments, assumptions or assertions that appear in various journals, magazines, monographs and reports.” JBI is clear that a single expert opinion cannot be used to represent the view of a group and explicit note appraisal checklist.

Another valuable initiative was launched by a group of health researchers who understand that decision makers are interested not only in whether an intervention works but also how it works and its components. They recognised the lack of tools for the critical appraisal of programme descriptions, descriptions of the implementation of interventions or programmes, such as in programme evaluation reports and other largely descriptive types of information. The group proposed a new approach to assessing the limitations of this ‘non-conventional’ evidence sources, the ‘Ways of Evaluating Important and Relevant Data’ (WEIRD) tool. The tool, comprised on instructions and a seven-page table with questions, guides the critical appraisal of sources not typically included in evidence reviews or evidence syntheses. These source materials are “not the product of a research process but may be generated as part of the routine planning and implementation of interventions, programmes or policies.” (Lewin, et al., 2019) The authors see WEIRD assessments as feeding into an assessment of how much confidence to place in findings from a synthesis.

Defining a more inclusive range of evidence

We found no use of the term “inclusive evidence” and only one-third of the initiatives made reference to grey literature. Further probing found multiple definitions of grey literature. The common dichotomy of published vs unpublished or informally published is not very enlightening as locally generated or contextualised evidence on practice may be published but not through commercial or academic channels. Defining informal publication outlets is also difficult and can reenforce discrimination.

Another term found was practice-derived evidence which was contrasted with rigorous research design. This also seems an unhelpful dichotomy. Quite a bit of academic education research focuses on policy and practice in action. Despite some of these distinctions being unhelpful, being aware of terms being used by different organisations will assist in defining what we mean as “more inclusive evidence”.

“Grey literature” is a challenge to identify and access. It is often not well represented in indexing databases, although academic libraries are growing their collections. It is surprising to find multiple definitions. The “Luxembourg definition” is a widely accepted definition in the scholarly community for grey literature, including OpenGrey, Karolinska Institute, Exeter University, and Monash University:

"information produced on all levels of government, academia, business and industry in electronic and print formats not controlled by commercial publishing" i.e. where publishing is not the primary activity of the producing body." Third International Conference on Grey Literature in 1997 (ICGL Luxembourg definition, 1997 - Expanded in New York, 2004).

This definition, that grey literature includes everything published by an entity whose primary activity is not publishing for profit (commercial), seems arbitrary and out of step with contemporary publishing. The Education field has many non-profit organisations that publish, complete with ISBN numbers, but whose main activity is training or research rather than publishing. Other characteristics might better define grey literature: anything unpublished or published informally, if “informally” that can be defined.

University College London defines it as “content that is produced and published by non-commercial private or public entities, including pressure groups, charities, and organisations such as the OECD, World Bank and WHO.” This would include publications with an ISBN, which for many defines a formal publication. Perhaps the most inclusive and simplest definition for grey literature is everything that is not a book or journal article.

Education.org has been using the definition of grey literature as “materials and research published specifically outside of the traditional commercial, academic publishing, and distribution channels.” (Paperpile 2020)3 This is recommended and is most useful if accompanied by a comprehensive list of examples of what constitutes grey literature. The types of grey literature appropriate for the analytic purpose should be decided ahead of the search. A more extensive list of types of grey literature is proposed here.

Since this definition encompasses a huge variety of evidence, it is helpful to have a list of types of sources that are included. These types may be useful for searching, classifying and appraising.

Figure 2: Examples of grey literature

Bibliographies

Pamphlets

Blogs

Patents

Clinical trials

Policy statements

Company Information

Posters

Conference papers/proceedings

Pre-print articles

Datasets

Presentations

Discussion Forums

Press releases

Dissertations and theses

Project implementation reports

Email discussion lists

Political declarations

Evaluations

Research reports

Government documents and reports

State-of-the art reports

Guidelines

Statistical Reports

Interview transcripts

Survey results

Legislation

Technical reports

Manuals

Technical specifications and standards

Market reports

Translations that are not commercial

Memoranda

Tweets

Minutes of meetings

Website pages

Newsletters

Wikis

Op-eds and letters to the editor

Working papers

3.4 Lessons for identifying and collecting inclusive evidence

Addressing bias at the start

Just as each piece of evidence may have a bias, the process of reviewing evidence is also subject to bias. This is as true for the review of a few pieces of evidence by a government official or of a large body of evidence by an analyst conducting an evidence synthesis. Bias can be introduced at any stage in the review process: formulating the review question; establishing what evidence will be included and excluded; finding the evidence; appraising the selected resources; and choosing what findings to make public. Bias distorts the results of a review, leads to false conclusions and is potentially misleading.  Bias may occur consciously or unconsciously on the part of the analyst when designing the review or interpreting the results. The effect is the production of incorrect conclusions that favour the researcher's beliefs, expectations, or commitments.

To avoid these problems in evidence reviews and syntheses, plans must be set out in advance in evidence review or synthesis protocols. In this protocol, analysts make clear what measures will be taken to reduce biases and the effects of the play of chance during the review process.

A similar process can be used in a country-level attempt to integrate locally generated evidence, in efforts to identify, collect, and ultimately appraise inclusive evidence for analysis in. Recommended rubrics, checklists, or review protocols can be developed to support at each stage.

Relevance of the evidence to the topic and end users is key to identifying evidence. The Sutherland and Wordley (2018) proposal for a subject-wide evidence synthesis suggests several channels and practices for identifying and collecting a wider range of evidence. These can be applied to any review, not only to an evidence synthesis, but are targeted at conventional publications.

  • Ask experts and compile list of intervention likely to be relevant,
  • Source widely, including every issue of subject journals instead of search terms,
  • Assess relevance of source based on title, abstract or the entire paper if necessary,
  • Only relevant sources are extracted, tagged, and stored.

What insights does Citizen Science hold?

Citizen Science (CS) is the practice of involving members of the public in a voluntary capacity in research projects, usually by involving them in collecting data (Irwin 2018). Less commonly, volunteer citizens are engaged in the interpretation of data, such as classifying image content (CitizenScience.gov). CS has been in use for decades, but widespread internet access boosted its use and smartphones increased the means to involve non-researchers in studies.

Potential benefits of CS

The benefits of tapping the non-professional workforce through CS goes beyond cost-saving. The engagement of citizens exposes them to science and helps to develop greater ownership of a field, for example working on environmental research projects can help develop a stewardship ethic. As Gollan (2013) argues, CS can “create a more scientifically literate society; building the capacity for people to take information they receive in their everyday lives and then being able to make informed choices based on…what they have learned.” In efforts to democratise evidence, a CS approach to evidence collection would also help validate the experiences of a wide range of education stakeholders.

The Use of CS

While the main fields using CS are biology, ecology, and astronomy, CS is described as a flexible concept which can be adapted and applied across diverse disciplines and situations. Most commonly, CS is used for data collection in scientist-led research projects that are designed from the top down. A general practice is outlined in Figure 1 below. CS is mostly used to collect quantitative data but there are examples of qualitative data collection, such as citizens without an academic background conducting interviews. CS is recognised in the literature as having potential use in community-oriented (bottom up) projects, for example where community members raise research questions, play a role in defining the focus, or analyse evidence.

For the purposes of democratising evidence, CS can be employed at a few stages, including contributing to topic selection, mapping a topic’s issue tree (scoping), and collecting sources of evidence. No examples of these uses were found. While there is no direct link between CS and accessing grey literature, crowdsourcing can be used in calls for existing sources of evidence, including grey literature. This was done with the Education.org prototype AEP evidence synthesis.

Quality Control in CS

Understandably, concerns have been raised about the quality of crowdsourced data in CS practice, with scientists themselves being suspicious of citizen-collected data for scientific purposes (Gollan, 2013) (Sharma, 2019; Santos-Fernandez & Mengersen, 2021). The concerns relate to biases that may be introduced through the use of volunteers, such as their vested interests in the data showing a particular outcome and their lack of preparedness to follow the data collection protocol. Biased inferences based on CS generated data can also result if the sampling effort is unknown and varying. To produce reliable evidence, the sampling process behind the data generated by volunteers must be considered (Sicacha Parada, 2020) (Sicacha Parada et al., 2021). A recognised example is a weekend bias in citizen science data reporting, since volunteers tend to have more time available on weekends compared to weekdays (Courter et. al, 2013). However, studies in the ecology field document an accuracy higher than 95% for data from volunteers (Kart, 2021; Gollan, 2013). In addition, some experts argue that trained scientists will produce disturbing rates of falsification (Sharma, 2019). Well-designed research projects and robust protocols for data gathering must be used to minimise that issue. Courter et al. argue that if known biases of CS data reporting are identified and addressed, the potential benefits can still be substantial.

Although there is no internationally recognized definition of CS, several institutions engaged in CS have published guidelines and handbooks to address quality assurance and documentation in CS (i.e., the European Citizen Science Association (ECSA, 2015) (ECSA, 2020) and the US Environmental Protection Agency (EPA, 2019a, EPA, 2019b). These recommendations focus on methodology, training participants, and templates for reliable documentation.

Body image

Figure: Generic Project Plan for CS by the US Environmental Protection Agency (EPA)

The following flow chart of a generic project outline is derived from EPA’s recommendations in order to assure data quality (USG EPA, 2021, P. 7 – 11).

Key takeaways on Citizen Science

Although not commonly used in the education field, Citizen Science principles and practices can be applied to our efforts to make evidence more inclusive:

  • Engage volunteers at the time of design to help shape the issues tree on the topic,
  • Solicit sources from, using a template or clear crowdsourcing call for sources,
  • Address quality control and minimise bias by designing and documenting a protocol and templates for reliable documentation. Be transparent by making these available to the public.

3.5 Lessons for classifying evidence, especially “grey literature”

It is helpful to tag studies prior to assessing their quality. However, as the table below indicates, most of the proposed labels are for research studies and do not include a broader range of evidence useful for decisionmakers, e.g. non-research evidence.

Type of research, research design and method (BE2 2015) 

Type of research 

Research design 

Typical analytic methods 

Primary and empirical studies 

Observational/descriptive or non-experimental research designs  

  • Cross-sectional regression analysis/large-n survey regression analysis (Quantitative) 
  • Cohort/longitudinal/panel data regression analysis (Quantitative) 
  • Analysis of interviews/focus group data (Qualitative) 
  • Analysis of ethnographic research (Qualitative) 
  • Case study research analysis (either) 
  • Political economy analysis (Qualitative or mixed) 
  • Mixed-methods research 

Quasi-experimental research designs  

  • Propensity score matching (Qualitative) 
  • Double difference methods (Quantitative, possibly with qualitative) 
  • Regression discontinuity designs (Quantitative) 

Experimental research designs (RCTs) 

  • Difference in difference  

Secondary research 

Systematic reviews  

Rigorous reviews  

Non-systematic reviews 

Non-systematic review (NSR) designs:  

  • Evidence papers  
  • Literature reviews  
  • Rapid reviews  
  • Policy analyses 

Theoretical or conceptual  

Rather, the list of types of grey literature and the Joanna Briggs Institute guidance suggest many more categories that can be added. 

3.6 Lessons for appraising the quality of evidence especially “grey literature” 

Principles 

The Joanna Briggs Institute appraisal checklist (Annex 2) and the WEIRD tool were already mentioned as guidance for appraising the quality of text and opinion sources as well as descriptive materials from programmes or interventions. The following table captures principles for appraising individual pieces of research. They are taken from two BE2 Guidance Notes, the second one specifically for qualitative research. Modifications need to be made to adapt these principles to appraise “grey literature”, especially evidence that is not research. 

Figure: Two examples of principles for appraising quality of research 

BE2 2015 Principles of High Quality Research 

DeJaeghere, et all 2020 (BE2 Guidance Note) 

Conceptual Framing  

Conceptual Framing 

Openness and Transparency 

 

Robustness of Methodology 

Robustness of Methodology 

Cultural Appropriateness/Sensitivity 

Culturally appropriate tools and analysis 

Validity  

 

 

Credibility (Sampling, pilot protocols, training) 

Reliability 

Reliability 

Cogency  

Openness and transparency; cogency 

See Annex 3. DeJaeghere, et al., 2020 for a checklist to assess the quality of qualitative research.

Reliability and Validity debates for qualitative research 

Some qualitative researchers (Lincoln and Guba 1985, p. 290) argue for abandoning the concepts of reliability and validity for qualitative research, and instead use “trustworthiness.” In DeJaeghere, et al., 2020, the authors reframe reliability and validity as dimensions of both qualitative and quantitative research (largely to be consistent with 2015 BE2 publication, Assessing the Strengthen of Education Evidence) to be addressed together through several techniques such as bias checks, triangulation, thick description, audit trail, and peer checking.  

 Use of scales 

A number of initiatives use a scale to indicate the overall strength of either a piece of evidence, or more commonly, the body of evidence on a specific issue. It is generally understood that the assignment of a particular score or grade is a judgement on the part of a reviewer. 

Body image

Figure: BE2 2015 proposes five levels for assessing the quality of an individual study.

4. Implications for IWG

4.1 Key recommendations

Decisionmakers deserve access to the best available evidence. But why is it important to make a wide range of evidence available to education decisionmakers? First, the culture of evidence use for decision-making in education is non-existent in some settings, and not as strong as needs to be. Secondly, worldwide, billions of dollars have been expended on education, yet only a tiny fraction of those investments were informed by evidence. In too many places, we are still designing and implementing the same policies spoken about fifty years ago. The evidence base available to decisionmakers is severely limited by the failure to tap a huge volume of work by NGOs, CSOs, and researchers, which includes contextualised, practice-derived evidence of great value. It is comprised of research and non-research evidence, analytic and descriptive. It is not tapped because, it is difficult to access and is not valued as evidence. This lack of value is in part because it is not classified and protocols for appraising its quality are not well developed or widely used. In the end, the use of data for decision-making can transform educational decision-making into a more objective, less politicized sector, resulting in a more cost-efficient use of limited resources and better learning outcomes.

While most evidence is difficult for decisionmakers to access and use. Increasing the use of education evidence in decisions regarding policy, planning, and practice requires that more contextualised and relevant evidence is available to decisionmakers. Evidence that is not available through traditional publication channels is particularly difficult to identify, access (collect), and appraise. This large volume of evidence, this non-conventional evidence, includes valuable insights from policy implementation and educational practice as well as a much wider range of voices and perspectives. It may be research that was not published through academic or commercial channels, but it might also not be research at all. It may be descriptive, or analytic, or first-hand accounts. The locally generated evidence is of particular interest to national decisionmakers because it contextualised and captures the experience of citizens and frontline education workers.

The following recommendations emerge from the landscaping exercise conducted:

  1. A greater focus is needed to increase the identification and collection of education evidence not published through traditional channels. This effort should draw on the growing interest in grey literature repositories, citizen science, and open access to published and unpublished sources.
  2. Create a classification system for tagging many types of evidence, drawing on recent Joanna Briggs Institute work on text and opinion (McArthur, et al., 2020) and guidance for qualitative research (DeJaeghere, et al., 2020) and types of “grey literature” types. This is needed to mitigate bias and support transparency, both of which are critical to increasing the trustworthiness of any analysis.
  3. Develop a rigorous method for appraising unpublished or informally published evidence, including evidence that is not research, drawing on such work as WEIRD (Lewin, et al., 2019), the Joanna Briggs Institute (McArthur, et al., 2020, and DeJaeghere, et al. (2020).
  4. At every stage, elevate context and relevance as features of evidence quality.
  5. Promote the documenting of process and make the list of sources public. This helps to amplify the work on NGOs, CSOs, and researchers and highlights the value of such evidence.

These recommendations require the engagement of many, at country, regional, and global levels. The convening of a technical group of stakeholders, an International Working Group, should bring alternative perspectives on education evidence to the task. Such a group can deliberate, develop and vet guidance for advancing these recommendations.

The International Working Group (IWG) should be conceptualised as both a bridge-building and democratising effort, bringing together a wide range of stakeholders in the evidence ecosystems (producers, synthesisers, users, practitioners, funders, etc.) to help expand the types of evidence, and therefore voices, available to decisionmakers beyond those published through academic and commercial channels.

It is proposed that the IWG will focus on the following activities:

  1. Develop a statement that articulates the rationale for making evidence available to decisionmakers more expansive. This relates to expanding the voices and perspectives represented in evidence for decisionmakers, broadening the range of sources accepted as evidence, including sources that are not research, such as descriptive texts, first-hand accounts, and practice-derived writing. This work will help democratise evidence and build bridges between knowledge actors, policymakers, and practitioners.
  2. Identify ways to facilitate the identification and collection of evidence that is not easily accessible. Hear from a wide range of knowledge producers and practitioners; debate terms such as grey literature, searching for how to explain the kinds of evidence most clearly; agree on a comprehensive list of types of sources to consider (perhaps starting with the grey literature list in the landscaping report) as an aid to searching and classifying.
  3. Propose a classification of grey literature which builds on but expands existing ones for conventional research, especially qualitative research.
  4. Develop guidance to appraise non-conventional evidence, building on but modifying the principles used for appraising research. The relevance of the evidence for the question and end users must be paramount.
  5. Learn from on the trial of this guidance in two countries and revise the guidance accordingly.
  6. Promote the use of IWG in home institutions and with partners.
  7. Monitor progress in making locally generated evidence more accessible. The LEARRN Results Framework will serve as the AMELP framework. Annual monitoring for five or more years is expected.

5. Broader lessons for Education.org Evidence Syntheses

5.1 Improve rigor and transparency overall

Strengthen how the transparency principle is put in practice. Documentation is essential for our efforts to develop an established standard methodology which we hope will become a gold standard for evidence syntheses serving decision makers. This includes:

  • A written protocol to identify and source evidence sources ahead of time, classifying/tagging by type, and criteria for appraising sources.
  • We then need to document the process followed along the way and have an explicit statement on dealing with bias. 
  • Use the more comprehensive list of types of grey literature and tag to allow a count and data presentation.
  • Publish the protocol.

5.2 Lessons for the Scope and Plan stage

  • Create means to consult with stakeholders on the issues tree. Ensure multiple perspectives, including end users, the education decisonmakers.

5.3 Lessons for the Identify and Catalog stage

  • Tap grey literature electronic databases more systematically. A more comprehensive set was compiled, including 31 containing grey literature, such as CODESRIA's Theses and Dissertations Catalogue, African Education Research Database (Annex 3). A more focused list of databases needs to be created, especially databases with grey literature.
  • Draw on the Citizen Science lessons for crowdsourcing evidence.
  • Consider the approach of Sutherland and Wordley (2018) who provide a strong argument for using a different approach to systematic reviews in fields—including international development and education—in which data are sparse or patchily distributed, or where studies vary greatly in design and generalizability. These situations need “a large-scale, cost-effective way of rigorously appraising information for applied fields.” (p. 365). Sutherland and Wordley (2018) recommend subject-wide evidence synthesis combines elements of systematic reviewing and mapping as well as expert assessment. The steps are as follows:
  1. Ask experts and compile list of intervention likely to be relevant
  2. Source widely, including every issue of subject journals instead of search terms
  3. Assess relevance of source based on title, abstract or entire paper if need be
  4. Only relevant sources are extracted, tagged, and stored
  5. Summarize key fundings of all studies
  6. Summarize each source, noting design, sample sizes, location in 1 para
  7. Obtain an expert assessment on intervention, addressing bias through multiple rounds of scoring and large team of assessors and an anonymous Delphi technique.
  8. Collate all the info in a synopsis.

While Sutherland and Wordley are not referring to grey literature, the above steps can be applied to an evidence synthesis that includes grey literature.

5.4 Lessons for the Screen and Tag stage (includes appraisal)

  • Assess the quality of secondary studies (evidence reviews, etc.) as well as individual piece of evidence. See text box below for example.

BE2 2015 proposes that bodies of evidence should be summarised in terms of four characteristics:  

  1. The (technical) quality of the studies constituting the body of evidence;  
  2. The size of the body of evidence (large, medium, or small and state number) 
  3. The context in which the evidence is set (global or context-specific);  
  4. The consistency of the findings produced by studies constituting the body of evidence (identical or similar; inconsistent with range of conclusions, possible contradictory). 

The following scale is proposed by BE2 2015 for assessing the quality of the body of evidence. It rejects a hierarchy of evidence and goes further to “suggest that bodies of evidence constituted by a diverse range of robust designs and methods (typically experimental and observational and both quantitative and qualitative in nature) are likely to be stronger than those that are reliant upon just one design, or just one or two methods.” (BE2 2015: 32). 

The four features can then be summarised as Strong, Medium, or Weak (BE2015: 34). 

5.5 Lessons for the Distil and Synthesis stage

  • Discuss and agree on how we use interviews in future evidence syntheses. It should be for clarification or to validate an interpretation rather than collect new information. We must be careful about going into primary research because this is a never-ending exercise and “a little bit of it” can bias our findings. Also, we have been explicit that this is not our focus. This distinction may not be easy to navigate.  
  • Considering using evidence maps for visual presentation in an evidence synthesis. FCDO typically expects that recommendations from syntheses of evidence will be summarised and visually represented through evidence maps and evidence briefs.
  • Indicate when data exists but is not made accessible. A gentle “name and shame” serves as part of the advocacy for open access. 

5.6 Lessons for Surface and Engage stage

The SURE Collaboration Guide 8. Informing and engaging stakeholders (SURE 2011) is focused on policy brief preparation and use. However, it is a useful framework for stakeholder engagement for every evidence synthesis, at several stages. Involvement starts prioritising topics, for example using the Global Council members. But it is critically important to consult and involve a range of various stakeholders in subsequent stages:

  • Clarifying the problem (Scope and Plan)
  • Deciding on and describing policy options (Scope and Plan)
  • Identifying and addressing barrier to implementation policy options (Scope and Plan)
  • Clarifying uncertainties and needs for M&E (Distil and Synthesise)
  • Organising and running policy dialogues (Surface and Engage)
  • Informing and engagement stakeholders in using a policy brief (Surface and Engage)

This suggests a stakeholder engagement plan from the start. The SURE guide provides questions to inform who to involve, to do what, how to choose, and types of involvement along a continuum and what potential outcomes (including capacity development) can result.

Related content