Literature DB >> 24303242

Characterization of the biomedical query mediation process.

Gregory W Hruby1, Mary Regina Boland, James J Cimino, Junfeng Gao, Adam B Wilcox, Julia Hirschberg, Chunhua Weng.   

Abstract

To most medical researchers, databases are obscure black boxes. Query analysts are often indispensable guides aiding researchers to perform mediated data queries. However, this approach does not scale up and is time-consuming and expensive. We analyzed query mediation dialogues to inform future designs of intelligent query mediation systems. Thirty-one mediated query sessions for 22 research projects were recorded and transcribed. We analyzed 10 of these to develop an annotation schema for dialogue acts through iterative refinement. Three coders independently annotated all 3160 dialogue acts. We assessed the inter-rater agreement and resolved disagreement by group consensus. This study contributes early knowledge of the query negotiation space for medical research. We conclude that research data query formulation is not a straightforward translation from researcher data needs to database queries, but rather iterative, process-oriented needs assessment and refinement.

Entities:  

Year:  2013        PMID: 24303242      PMCID: PMC3845777     

Source DB:  PubMed          Journal:  AMIA Jt Summits Transl Sci Proc


Introduction

The accelerated adoption of electronic health records (EHRs) systems nationwide, fueled in part by administrative initiatives, has made an unprecedented amount of data available. 1 , 2 Appropriate reuse of such “big data” from healthcare processes is indispensible for achieving comparative effectiveness research (CER). 3 Expanding data access for clinical and translational researchers has long been an important priority for accelerating clinical and translational research. Many institutions employ query analysts to translate data requests from medical researchers into executable database queries. Unfortunately, this approach is not scalable, partly because many requests contain vague or nonspecific concepts that are hard to understand for query analysts, who usually have limited medical domain knowledge. Typically, query analysts have to consult researchers many times via emails or phone calls to clarify the query details. To relieve the burden on query analysts, a variety of data query tools were developed. 4 – 8 Notable ones include Informatics for Integrating Biology and Bedside (i2b2), 9 – 11 the Visual Aggregator and Explorer (VISAGE), 12 and the Stanford Translational Research Integrated Database Environment (STRIDE). 13 I2b2 enables users to drag and drop concepts to construct queries. The modified Web version, SHRINE, also enables federated queries across multiple databases. 9 Similarly, VISAGE is an ontology-driven visual query interface that recommends concepts for query formulation. Such tools generally require users to specify or select concepts for query formulation, which can be a significant challenge for researchers who usually have limited knowledge of the organization and coding of the data or the sensitivity and specificity of terms in the database. This problem becomes worse as databases increase in size and complexity. Ideally, researchers should interrogate databases autonomously. One approach to achieve this goal is to support computer-based reference interviews with biomedical researchers. “A reference interview is a conversation between a librarian and a library user, usually at a reference desk, in which the librarian responds to the user’s initial explanation of his or her information need by first attempting to clarify that need and then by directing the user to appropriate information resources”. 14 Query analysts, like librarians, often use a negotiation process to comprehend the needs of the researcher. 15 , 16 However, at this point, little is known about common steps and their temporal relationships during the biomedical query mediation process. Query analysts often do not have a reference interview template to guide them through the query mediation process. Therefore, this study reports our analysis of the query mediation dialogues between a query analyst and medical researchers and our findings of the characteristics of the biomedical query mediation process. This study extends the work of a previous poster presented at the 2012 AMIA Fall Symposium entitled, “Analysis of Query Negotiation between a Researcher and a Query Expert.” 17

Data and Methods

Data:

Between July 2011 and January 2012, we recorded and transcribed 31 discussions for 22 medical research projects between one query expert (QE) and eight medical researchers (MRs) at the Columbia University Department of Urology. The Columbia University Medical Center Institutional Review Board approved this study ( IRB-AAAJ8850 ). Figure 1 shows 5 example dialogue acts. In the context of this paper, a dialogue act is one exchange of speech. We arrived at 3160 dialog acts for the 31 query mediation sessions.
Figure 1.

Example Dialogue Acts

Annotation Schema Development:

We used the dialogue acts from 10 randomly selected projects to develop a dialogue act classification schema. We first derived the common tasks of dialogue acts, such as understanding the clinical process, identifying available data, and explaining data characteristics. Then, we grouped the tasks by their corresponding aspect of the query mediation process, such as stages of mediation, data request complexity, and interpretation of requester response. We decided to classify dialogue acts along the “Stages of Mediation” aspect in order to see the temporal patterns of dialogue acts. We iteratively designed and tested a classification schema on sample transcripts and finalized the schema with group consensus among three independent raters.

Dialogue Act Annotation:

Three raters (GH, MB, JG) independently annotated all the 3160 dialogue acts. Each dialogue act was annotated with at least one classification code. We assessed inter-rater agreement with the kappa statistic. For dialogue acts with inter-rater disagreement, we reached consensus by accepting the pair-wise consensus between GH and MB first, GH and JG next, and MB and JG last. We resolved the remaining disagreements of the clinical content of dialogue acts with GH codes.

Data Analysis:

We used the consensus annotation results for further dialogue flow analysis. We normalized the query negotiation space for the 22 projects to the median number of dialogue acts by either condensing or expanding the conversation sets for the 22 projects. We aggregated the annotated content of the 22 projects into one representation of the negotiation space. We used descriptive statistics and graphs to visualize this space.

Results

A Dialogue Act Classification Schema for Mediate Query Conversations

The minimum, median, and maximum numbers of dialogue acts in a project were 27, 134, 323, respectively. The tasks we identified corresponding to the aspect of “Query Mediation Steps” are (1) State the Problem, (2) Locate Data Elements in EHRs, (3) Project Re-Iteration, (4) Discuss Study Design, and (5) Confirm Completed Process. These served as the basis for our coding book ( Figure 2 ). Tasks were iteratively organized into a hierarchical structure ( Figure 2 ) to be used to describe the dialogue acts of the mediation process between the QE and MR. Figure 2 shows the coding schema we used to annotate the dialogue acts of the query negotiation space. Our inter-rater kappa score over all the dialogue acts was 0.61.
Figure 2.

The Classification schema for Dialogue Acts in Query Mediation

Temporal Distribution of Dialogue Act Classes

Figure 3 illustrates the broad variety of issues discussed between a QE and MR. This figure also represent the aggregate of all 22 projects into one normalized space. The y-axis represents the total number of codes used to annotate a particular conversation act defined by the x-axis. For example, 62 codes were used to annotate the first conversation act of all 22 projects. Throughout the conversation, the majority of the discussion surrounds the clinical process. However, as the conversation concludes, greater attention is drawn toward the research workflow clarification. Additionally, as the conversation concludes, the QE and the MR discuss IRB and privacy policy.
Figure 3.

Temporal Distribution of Dialogue Acts in a Normalized Mediated Query Conversation Session

A closer look on the Discussion of the Clinical Process

Figure 4 shows how the clinical content of the space is left-skewed towards the beginning of the conversation and trails off at the end. The blue thin line represents the aggregated clinical variable codes, 2.0 (“Explain the Clinical Process”). The blue thick line represents the trend of this variable.
Figure 4.

Discussion of the Clinical Process over the Course of a Normalized Conversation Session

Temporal Distribution of Study Design and Research Workflow Discussions

Figure 5 shows two classes from the coding schema, 4.0 Discuss Study Design (blue line) and 5.0 Clarify Research Workflow (red line). The start of the conversation supports the development of the study design. The middle of the conversation exchanges these two classes cyclically until the end of the conversation, where research workflow emerges as the dominant class.
Figure 5.

Discussion about Study Design and Workflow Issues throughout a Normalized Conversation

Discussion

As the health record transforms and migrates to the electronic form, data requests for research purposes are likely to increase. The volume of these requests will quickly overwhelm those responsible for querying these data. Non-mediated means for data queries exist but fail to fully satisfy researcher’s data needs. Instead, a mediated data extraction process is needed. However, little is known about the negotiation space between the QE and the MR. Zhang et al. briefly describe this process in their “data access paradigm model”. 7 We were able to identify several classes that fall under “stages of the negotiation process.” After several iterations and reductions to the class list, the granularity of the classes was expanded to create our annotation schema for dialogue acts for mediated queries. Although, we had an inter-rater kappa score of 0.61, we do not expect for this coding book to generalize to all other research query mediation processes, but rather to describe the content of this specific negotiation space. This coding schema will allow us to study the progression of conversation and inform the design of a structured interview between QE and MR. The initial illustration of the negotiation space ( Figure 3 ) is a clear representation of the complexity that exists. Furthermore, we interpret this result as a clear refutation of the idea that data needs assessment is a simple and easy process. A significant amount of QE and MR investment is needed to reach an understanding of what the data needs are for any given project. This represents a critical part of the process that occurs in order for a consensus to be reached regarding the researcher’s data needs. We interpret the clinical content illustration ( Figure 4 ) to represent a clear presentation of potential clinical variables presented by the researcher. Difficult clinical concepts, discussed over the course of the conversation, are explored until an understanding is reached and the clinical content drops off toward the end of the conversation space. Figure 5 provides insight regarding how a conversation reaches consensus. It shows how a conversation moves from a theoretical description of data elements to a practical project management discussion. Of particular interest is the middle of the conversation space, where an iterative exchange is occurring between these two classes (Study Design and Research Workflow) of dialogue acts.

Limitations

This study contains two major limitations. First, we only analyzed the conversation space of one QE (GH) with medical researchers from one academic department. Furthermore, this QE was intensively involved with the department’s research program. The QE facilitated not just data access but also study design and project management. As such, the conversation space may cover more issues then traditional query negotiations that exist between other QE and MR.

Conclusion

To the best of our knowledge, this study represents the first attempt to understand the mediated query dialogues between a query expert and a medical researcher. Our results confirmed that the query negotiation space is not a straightforward translation of a researcher’s needs, but rather an iterative process necessary to reach an understanding of what those needs are. Query mediation represents a process-based needs assessment and clarification. The results of this study prepare us for our next steps, which are to extract common dialogue elements in mediated query processes and to model the conversation flow in order to inform the design of structured query negotiation, towards the development of an intelligent virtual medical data librarian.
  12 in total

1.  The Electronic Data Methods (EDM) forum for comparative effectiveness research (CER).

Authors:  Erin Holve; Courtney Segal; Marianne Hamilton Lopez; Alison Rein; Beth H Johnson
Journal:  Med Care       Date:  2012-07       Impact factor: 2.983

2.  STRIDE--An integrated standards-based translational research informatics platform.

Authors:  Henry J Lowe; Todd A Ferris; Penni M Hernandez; Susan C Weber
Journal:  AMIA Annu Symp Proc       Date:  2009-11-14

3.  Architecture of the open-source clinical research chart from Informatics for Integrating Biology and the Bedside.

Authors:  Shawn N Murphy; Michael Mendis; Kristel Hackett; Rajesh Kuttan; Wensong Pan; Lori C Phillips; Vivian Gainer; David Berkowicz; John P Glaser; Isaac Kohane; Henry C Chueh
Journal:  AMIA Annu Symp Proc       Date:  2007-10-11

4.  The Shared Health Research Information Network (SHRINE): a prototype federated query tool for clinical data repositories.

Authors:  Griffin M Weber; Shawn N Murphy; Andrew J McMurry; Douglas Macfadden; Daniel J Nigrin; Susanne Churchill; Isaac S Kohane
Journal:  J Am Med Inform Assoc       Date:  2009-06-30       Impact factor: 4.497

5.  Access to data: comparing AccessMed with Query by Review.

Authors:  G Hripcsak; B Allen; J J Cimino; R Lee
Journal:  J Am Med Inform Assoc       Date:  1996 Jul-Aug       Impact factor: 4.497

6.  Serving the enterprise and beyond with informatics for integrating biology and the bedside (i2b2).

Authors:  Shawn N Murphy; Griffin Weber; Michael Mendis; Vivian Gainer; Henry C Chueh; Susanne Churchill; Isaac Kohane
Journal:  J Am Med Inform Assoc       Date:  2010 Mar-Apr       Impact factor: 4.497

7.  Electronic medical records for genetic research: results of the eMERGE consortium.

Authors:  Abel N Kho; Jennifer A Pacheco; Peggy L Peissig; Luke Rasmussen; Katherine M Newton; Noah Weston; Paul K Crane; Jyotishman Pathak; Christopher G Chute; Suzette J Bielinski; Iftikhar J Kullo; Rongling Li; Teri A Manolio; Rex L Chisholm; Joshua C Denny
Journal:  Sci Transl Med       Date:  2011-04-20       Impact factor: 17.956

8.  Implementation of a deidentified federated data network for population-based cohort discovery.

Authors:  Nicholas Anderson; Aaron Abend; Aaron Mandel; Estella Geraghty; Davera Gabriel; Rob Wynden; Michael Kamerick; Kent Anderson; Julie Rainwater; Peter Tarczy-Hornoch
Journal:  J Am Med Inform Assoc       Date:  2011-08-26       Impact factor: 4.497

9.  VISAGE: A Query Interface for Clinical Research.

Authors:  Guo-Qiang Zhang; Trish Siegler; Paul Saxman; Neil Sandberg; Remo Mueller; Nathan Johnson; Dale Hunscher; Sivaram Arabandi
Journal:  Summit Transl Bioinform       Date:  2010-03-01

10.  Next-generation phenotyping of electronic health records.

Authors:  George Hripcsak; David J Albers
Journal:  J Am Med Inform Assoc       Date:  2012-09-06       Impact factor: 4.497

View more
  13 in total

Review 1.  Review and evaluation of electronic health records-driven phenotype algorithm authoring tools for clinical and translational research.

Authors:  Jie Xu; Luke V Rasmussen; Pamela L Shaw; Guoqian Jiang; Richard C Kiefer; Huan Mo; Jennifer A Pacheco; Peter Speltz; Qian Zhu; Joshua C Denny; Jyotishman Pathak; William K Thompson; Enid Montague
Journal:  J Am Med Inform Assoc       Date:  2015-07-29       Impact factor: 4.497

2.  What Is Asked in Clinical Data Request Forms? A Multi-site Thematic Analysis of Forms Towards Better Data Access Support.

Authors:  David A Hanauer; Gregory W Hruby; Daniel G Fort; Luke V Rasmussen; Eneida A Mendonça; Chunhua Weng
Journal:  AMIA Annu Symp Proc       Date:  2014-11-14

3.  A multi-site cognitive task analysis for biomedical query mediation.

Authors:  Gregory W Hruby; Luke V Rasmussen; David Hanauer; Vimla L Patel; James J Cimino; Chunhua Weng
Journal:  Int J Med Inform       Date:  2016-06-16       Impact factor: 4.046

4.  Clinical Research Informatics for Big Data and Precision Medicine.

Authors:  C Weng; M G Kahn
Journal:  Yearb Med Inform       Date:  2016-11-10

5.  Leveraging dialog systems research to assist biomedical researchers' interrogation of Big Clinical Data.

Authors:  Julia Hoxha; Chunhua Weng
Journal:  J Biomed Inform       Date:  2016-04-08       Impact factor: 6.317

6.  DREAM: Classification scheme for dialog acts in clinical research query mediation.

Authors:  Julia Hoxha; Praveen Chandar; Zhe He; James Cimino; David Hanauer; Chunhua Weng
Journal:  J Biomed Inform       Date:  2015-11-30       Impact factor: 6.317

Review 7.  Facilitating biomedical researchers' interrogation of electronic health record data: Ideas from outside of biomedical informatics.

Authors:  Gregory W Hruby; Konstantina Matsoukas; James J Cimino; Chunhua Weng
Journal:  J Biomed Inform       Date:  2016-03-10       Impact factor: 6.317

8.  A data-driven concept schema for defining clinical research data needs.

Authors:  Gregory W Hruby; Julia Hoxha; Praveen Chandar Ravichandran; Eneida A Mendonça; David A Hanauer; Chunhua Weng
Journal:  Int J Med Inform       Date:  2016-04-02       Impact factor: 4.046

9.  A pattern learning-based method for temporal expression extraction and normalization from multi-lingual heterogeneous clinical texts.

Authors:  Tianyong Hao; Xiaoyi Pan; Zhiying Gu; Yingying Qu; Heng Weng
Journal:  BMC Med Inform Decis Mak       Date:  2018-03-22       Impact factor: 2.796

10.  Research Data Explorer: Lessons Learned in Design and Development of Context-based Cohort Definition and Selection.

Authors:  Adam Wilcox; David Vawdrey; Chunhua Weng; Mark Velez; Suzanne Bakken
Journal:  AMIA Jt Summits Transl Sci Proc       Date:  2015-03-25
View more

北京卡尤迪生物科技股份有限公司 © 2022-2023.