| Literature DB >> 27768760 |
Stephen A Gallo1, Joanne H Sullivan1, Scott R Glisson1.
Abstract
Although the scientific peer review process is crucial to distributing research investments, little has been reported about the decision-making processes used by reviewers. One key attribute likely to be important for decision-making is reviewer expertise. Recent data from an experimental blinded review utilizing a direct measure of expertise has found that closer intellectual distances between applicant and reviewer lead to harsher evaluations, possibly suggesting that information is differentially sampled across subject-matter expertise levels and across information type (e.g. strengths or weaknesses). However, social and professional networks have been suggested to play a role in reviewer scoring. In an effort to test whether this result can be replicated in a real-world unblinded study utilizing self-assessed reviewer expertise, we conducted a retrospective multi-level regression analysis of 1,450 individual unblinded evaluations of 725 biomedical research funding applications by 1,044 reviewers. Despite the large variability in the scoring data, the results are largely confirmatory of work from blinded reviews, by which a linear relationship between reviewer expertise and their evaluations was observed-reviewers with higher levels of self-assessed expertise tended to be harsher in their evaluations. However, we also found that reviewer and applicant seniority could influence this relationship, suggesting social networks could have subtle influences on reviewer scoring. Overall, these results highlight the need to explore how reviewers utilize their expertise to gather and weight information from the application in making their evaluations.Entities:
Mesh:
Year: 2016 PMID: 27768760 PMCID: PMC5074495 DOI: 10.1371/journal.pone.0165147
Source DB: PubMed Journal: PLoS One ISSN: 1932-6203 Impact factor: 3.240
Definitions for Scientific Merit and Reviewer Expertise Scoring.
| 1.0–1.9 | EXCEPTIONAL: The scientific merit of the proposal probably places it in the top 10% of proposals in its area of research; it warrants the highest priority for support. This category should be used only for truly outstanding proposals. A score of 1 indicates a very high level of scientific merit. |
| 2.0–2.9 | GOOD: The scientific merit of the proposal is such that it warrants high priority for support. A score of 2 indicates a significant level of scientific merit. |
| 3.0–3.9 | FAIR: The scientific merit of the proposal is not impressive, and it is probable that it does not warrant support as submitted. If the topic of the proposal is of particular interest, partial support may be warranted. Full support is unlikely to be appropriate. A score of 3 indicates only a moderate level of scientific merit. |
| 4.0–4.9 | DEFICIENT: The scientific merit of the proposal is low. The proposal is flawed, and support is unlikely to be justifiable. A score of 4 indicates a low level of scientific merit. |
| 5.0 | REJECT: The proposal has very serious deficiencies; it should not be supported under any circumstances. A score of 5 indicates a rejection of the work by the reviewers. |
| 1.0–1.9 | The proposal is in your specific area of active research. Your knowledge of current publications is thorough. |
| 2.0–2.9 | The proposal is in your general area of active research. Your knowledge of the literature is reasonably current. You could apply the techniques of the proposal with little difficulty. You have some ongoing communication with workers in the area of the proposal. |
| 3.0–3.9 | The proposal is outside your general area of active research, but it is related. You have knowledge derived from interest in the major discipline embracing the specific proposal, but have little or no contact with other workers active in similar research. |
| 4.0–4.9 | The proposal is not related to your active interest and is no more than peripheral to your major discipline. |
| 5.0 | The proposal is not related to your major discipline, and your knowledge is only derived through supplemental reading and interest in general science. |
Fig 1Average SM Score per Application.
Average SM score for each application versus the rank order by average SM score, with error bars representing standard error (2009–2012).
Reviewer and applicant demographics (2009–2012).
| Reviewer Demographics (Total Reviewers = 1044) | Applicant Demographics (Total Proposals = 725) | |||
|---|---|---|---|---|
| Factors | N | % | N | % |
| Male | 799 | 77 | 619 | 85 |
| Female | 245 | 23 | 106 | 15 |
| Junior Position | 276 | 26 | 121 | 17 |
| Non-Junior Position | 768 | 74 | 604 | 83 |
| Academia | 965 | 92 | 470 | 65 |
| Non-Academia | 79 | 8 | 255 | 35 |
| No MD Degree | 716 | 69 | 474 | 65 |
| MD Degree | 328 | 31 | 251 | 35 |
Multi-level regression comparison of random-intercept models.
| Model 1 | Model 2 | Model 3 | Model 4 | Model 5 | |
|---|---|---|---|---|---|
| Baseline Across Applications | RE | RE + Reviewer Demographics | RE + Reviewer and Applicant Demographics | RE + Seniority + Sector | |
| Variance Across Proposals | 0.219 (0.042)* | 0.213 (0.042)* | 0.204 (0.042)* | 0.194 (0.042)* | 0.192 (0.043)* |
| Residual Variance | 0.740 (0.023)* | 0.722 (0.023)* | 0.717 (0.023)* | 0.714 (0.023)* | 0.716 (0.023)* |
| Intercept | 2.77 (0.03)* | 2.77 (0.03)* | 2.80 (0.04)* | 2.74 (0.05)* | 2.73 (0.04)* |
| Reviewer Expertise | -0.15 (0.02)* | -0.09 (0.04)* | -0.09 (0.05) | -0.09 (0.03)* | |
| Change in 2LL | 38.9* | 36.8* | 20.9* | 16.0* | 36.9* (compared to Model 2) |
| 0.026 | 0.051 | 0.070 | 0.075 | 0.075 |
Analysis based on z-score of RE. Asterisk indicates statistical significance (p<0.05). Standard error is reported in parentheses. Each model was compared to the previous model (unless noted otherwise) through the calculation of deviance, as measured by the change in -2 log likelihood. All main effects and interactions with RE were included. Model 1 was compared to a fixed intercept model.
Fig 2Scatterplot of SM versus RE.
Scatterplot and linear regression fit of SM versus RE data with gray area representing 95% confidence intervals (2009–2012).
Fig 3Reviewer Seniority Scatterplots of SM versus RE.
Scatterplot and linear fit of raw SM versus RE scoring data of all evaluations by junior reviewers (in red) and by senior reviewers (in blue). The shaded area represents 95% confidence interval.
Fig 4Applicant Seniority Scatterplots of SM versus RE.
Scatterplot and linear fit of raw SM versus RE scoring data of all evaluations of junior applicants (red) and non-junior applicants (blue). The shaded area represents 95% confidence interval.
Summary of Model 5 (RE + Seniority + Research Sector).
| Variance Across Proposals | 0.192 (0.043)* |
| Residual Variance | 0.716 (0.023)* |
| Intra-proposal Correlation | 0.21* |
| Inter-Rater Reliability | 0.35 |
| R2 | 0.075 |
| Intercept | 2.73 (0.04)* |
| Reviewer Expertise (RE) | -0.09 (0.03)* |
| Junior Reviewer (RevJ) | 0.05 (0.06) |
| Junior Applicant (PIJ) | 0.16 (0.07)* |
| Non Academic Reviewer (RevNonAc) | -0.48 (0.12)* |
| Non-Academic Applicant (PINonAc) | 0.07 (0.06) |
| Reviewer Expertise * Junior Reviewer | -0.15 (0.06)* |
| Reviewer Expertise * Junior Applicant | -0.16 (0.07)* |
| Non Academic Reviewer (RevNonAc) * Non-Academic Applicant (PINonAc) | 0.41 (0.18)* |
Model 5 (random intercept model including fixed effects from RE, RevJ, PIJ, PINonAc and RevNonAc) coefficient estimates are listed. Analysis is based on z-score of RE. Standard error is reported in parentheses and asterisks indicate statistical significance (p<0.05). Results are broken out by random and fixed components (including both main and RE interaction effects). In addition, estimates of intra-proposal correlation and inter-rater reliability are provided.
Reviewer scoring and fundability agreement.
| Agreement on Fundability (Higher Expertise) | Agreement on Fundability (Lower Expertise) | Average Score Difference (Higher Expertise) | Average Score Difference (Lower Expertise) | |
|---|---|---|---|---|
| 33% | 35% | 0.66 (0.05) | 0.57 (0.05) | |
| 82% | 81% | 1.09 (0.05) | 0.95 (0.04) |
Inter-reviewer agreement between two reviewers assigned the same application on fundability, based on a 2.0 funding threshold (less than 2.0 is arbitrarily deemed fundable). This is shown for fundable and unfundable applications and for higher and lower average reviewer expertise (high is higher than median RE of 1.65; low is lower than median). Also average scoring difference (absolute differences) between assigned reviewers is shown with a similar breakdown (standard error shown in parentheses).