The UK Terminology Centre Data Migration Workbench (DMWB) is designed to support the NHS Primary Care Summary Care Record, Primary Care Systems of Choice and Data Migration programs. This tool demonstrates some of the properties and advanced uses of the data migration and mapping products published by the UKTC and the terminologies and classifications that they link.
The workbench uses SNOMED CT to perform novel and sophisticated analyses of patient data. It has immediate 'off the shelf' international utility despite the inclusion of the UK-only terminologies within the standard tool distribution.
The software contains SNOMED CT, Read Codes Version 2 and CTV3, maps between these and maps to ICD-10 International Edition (UK map not the same as the SNOMED International one) and OPCS Classification of Interventions and Procedures (OPCS-4). The Workbench modules support:
Searching and browsing the hosted code systems;
Viewing maps between the hosted code systems;
Authoring analytics subsets (i.e. terminology query predicates) and their testing, maintenance and 'translation' between code systems;
Electronic Patient Record (EPR) data quality analysis and data repair; and
EPR reporting and case mix analysis.
The Queries Tool offers advanced functionality for authoring, maintaining and testing query code sets (called 'clusters') or subset definitions within any of the supported terminologies or classifications. One major application is to produce query sets which will return comparable results from patient records encoded with different code systems. To assist with this, the tool translates subset definitions expressed using one terminology into subset definitions expressed using another, allowing refinement by manual editing.
Electronic Patient Record Data
The EPR Data Tool provides an environment for loading and analyzing patient data, either to design query specifications or as part of data quality and case mix analytics. The analytics functions include detection and management of coding data quality issues such as:
Records with inactive SNOMED CT codes;
Records with codes from inappropriate SNOMED CT hierarchies e.g. a diagnosis recorded using a concept from the substance hierarchy (e.g. 419442005 |ethyl alcohol|) rather than a code from the disorder hierarchy (e.g. 25702006 |alcohol intoxication| or 7200002 |alcoholism|).
The tool also enables the rapid repair of such data by substituting inappropriate codes with more appropriate ones. This service is performed in an offline reporting environment.
The workbench data analytics tool runs cluster queries, combined with demographic data, to perform clinically valuable case finding, case mix and caseload analysis.
The 'Overview' Report module includes:
Basic demographics (population age, sex, ethnicity);
Analyses of episodes with a SNOMED CT code;
Counts by SNOMED CT supercategory;
The Trends module analyzes the frequency with which individual SNOMED CT codes are used in the EPR instance data, looking for those whose recording frequency has changed over the course of the data collection period.
The Induce module performs a more sophisticated analysis of case mix and caseload trends within a clinical department. Instead of returning the most frequently used individual codes, the Induce module attempts to identify the most frequently used types of codes. For example, an emergency department may use roughly 500 different SNOMED CT codes for a laceration in a particular anatomical location. While none of the site-specific codes may appear in a list of most common codes, the descendants of 312608009 | laceration| may collectively account for a significant part of the department's workload.
The Graphs tool performs fundamentally the same query and search operations, but generates graphs based on the patients or episodes identified, showing e.g. the age:sex distribution of patients in a defined casemix cohort, or the changing incidence of one or more specified clinical phenomena (e.g. disease presentation, or procedure performed) by year, quarter, month or day of the week. These graphs can be copied into documents.
List of the 15 most common SNOMED CT codes for each age cohort.
1
This section includes brief reviews of a variety of projects which implement or support analytics over SNOMED CT enabled data.
We welcome additional input to this section and anticipate updates to this report as new information becomes available.


Kaiser Permanente (KP) has been involved in the development of SNOMED CT since its inception. Preceding this, KP collaborated with the College of American Pathologists in the 1990's on the immediate predecessor of SNOMED CT (SNOMED-RT). Some of the very earliest deployments of SNOMED CT have been within KP electronic patient record systems.
The terminology services deployed within the KP HealthConnect electronic health record illustrate the practical use of SNOMED CT as a key reference terminology within a multi-coding system environment. KP is also at the forefront of realizing new possibilities offered by SNOMED CT using its description logic capabilities.
Convergent Medical Terminology (CMT) is KP's Enterprise Terminology System. While the KP HealthConnect EHR system is built by Epic (see case study Epic), the CMT is proprietary to Kaiser Permanente. CMT hosts several components:
Standard reference terminologies
End user terminology (e.g. the terms presented to clinicians or patients)
Administrative codes and classifications (e.g. ICD-9-CM, ICD-10-CM, CPT4, HCPCS)
Analytics services (querying and decision support)
Request submission for new terms
CMT uses SNOMED CT as a reference terminology, taking advantage of its poly-hierarchy and definitional attributes to support advanced analytics – for example:
Identifying patient cohorts with certain conditions for Population Care.
Identifying subsets for use as "input criteria" for KPHC decision support modules, such as Best Practice Alerts, Reminders, etc.
Performing queries such as "find all conditions where |causative agent| is |Aspergillus (organism)|"
In September 2010 Kaiser Permanente, IHTSDO and the US Department of Health and Human Services jointly announced KPs donation of their CMT content and related tooling to SNOMED International. The donation consists of terminology content (including several CMT subsets), tools to help create, manage and quality control terminology.
KP in collaboration with the Information Systems Group at Oxford University are investigating how to perform complex queries efficiently across extremely large numbers of patient records. The team at Oxford University has developed an open source triple store (i.e. 'subject-predicate-object') database called RDFox. RDFox is highly scalable and performant 'Not Only SQL' database readily distributed across parallel processing units. RDFox is an implementation of the W3C Resource Description Framework (RDF) standard, which supports OWL-RL description logic.
In this collaborative project, clinical data is being represented in OWL-RL as 'entity-role-act' triples. This uses a logical model (with Entities in Roles participating in Acts) that is similar to HL7 V3's Reference Information Model. OWL-RL and Datalog rule language is being used to reason over hundreds of millions of patient data triples. While SNOMED CT expressions cannot be fully represented in OWL-RL, RDFox performs the preliminary large-scale clinical data retrieval to return a far smaller record set. This smaller set is then processed using a richer featured but less performant description logic reasoner supporting SNOMED CT.
Prototype work has been completed using real patient data, including observations for diabetes (coded using SNOMED CT) and observations of Hemoglobin A1C levels. Datalog instructions and SPARQL queries were used to calculate Healthcare Effectiveness Data and Information Set quality measures for diabetes management – for example, both numerators and denominators for the Diabetes HgB A1C report.
Migration to native SNOMED CT electronic patient records is in progress in the United Kingdom National Health Service (NHS). In order to promote interoperability, usability and activity reporting, the NHS introduced a national standard set of imaging codes in 2005 – the National Clinical Imaging Procedure code set (NCIP).
While SNOMED CT was the prime candidate for populating the NCIP, many Radiology Information Systems (RIS) and Picture Archiving and Communication (PAC) systems at the time could not accommodate SNOMED CT 18-digit concept identifiers or (up to) 255 character descriptions without disruptive and costly software changes. There was also no consistent way to represent laterality of procedures, and some legacy systems required the creation of separate orderable items for each laterality – for example 'Plain X-ray left wrist', 'Plain X-ray right wrist', and 'Plan X-ray both wrists'. For these reasons, the NCIP code set was developed based on SNOMED CT, but with the addition of unique identifiers compatible with legacy system's character limitations (6 alphabetic characters), up to 40 character human readable descriptions, and additional laterality metadata. For example, is represented within NCIP as:
The National Release Centre of Denmark (National eHealth Authority) produces a SNOMED CT drug extension for medications. The Danish SNOMED CT drug extension was primed by data extraction, cleansing and conversion of content from the Danish Medicine Agency Database (DKMDB), which is primarily meant for pricing and stock handling. The DKMDB was then complemented with SNOMED CT substances and their unique IDs. The Danish SNOMED CT drug extension includes information such as trade names, substances, dose forms, strengths and units of measure.
Building upon the Danish drug extension, the National eHealth Authority is working to introduce centralized decision support (CDS) services for both primary care and hospital prescribing systems.
The CDS server will respond to web service requests from the various electronic medication systems and return alerts and other prescribing information
Allergies Register
A group of allergy specialists, family practitioners and CDS experts are developing a standard set of information to be used in a patient drug allergy register. A SNOMED CT subset, from the Drug Allergy (disorder) sub hierarchy in the Findings hierarchy, is used to document allergies. .
1
Laterality_ID
Laterality
Short_Code
Preferred
60027007
51440002
Right and left
XWRIB
XR Wrist Both
60027007
7771000
Left
XWRIL
XR Wrist Lt
60027007
24028007
Right
XWRIR
XR Wrist Rt
NCIP short codes are 'meaningful', in that the modality of the procedure is defined by the first character of the code, and the finding site and laterality are both explicitly represented in the code.
Each hospital submits mandatory data extracts using NCIP from both legacy and SNOMED CT capable RIS. In addition to details of the imaging procedures, information about the referral source, patient type, demographics and times of each imaging related event are also collected centrally. The data from all sites is then combined and multiple reports are extracted. Hospitals can view their activity data via the iView web based reporting tool and compare their activity with other centers.
Analytics on this central platform are wholly SNOMED CT based. SNOMED CT hierarchies support sophisticated reports – for example, the monthly waiting times for Magnetic Resonance Imaging excluding Cardiac MRI and MRI guided procedures is specified as:
Includes hierarchy << 113091000 | Magnetic resonance imaging|
Excludes hierarchy << 258177008 | Magnetic resonance imaging guidance|
Excludes hierarchy << 241620005 | Magnetic resonance imaging of heart|
1
SCT ID
Allergy alerts are enabled by the relationships in SNOMED CT between allergy disorders and substances (via the |causative agent| attribute), and relationships between drug products and substance concepts (via the |has active ingredient| attribute).
Based on an existing service, with data primarily drawn from peer-reviewed literature, the interaction database describes 2,500 interactions between different drugs based on their ingredients.
The database contains a short description of all interactions and a recommendation of how the physician can handle the interaction. The ingredients have been linked to SNOMED CT substances to directly inform the decision support service.
The risk situation database contains drugs evaluated by experts as being potentially dangerous in specific situations. Drug products, ingredients and dose forms are converted to SNOMED CT concepts, which thus contribute to the decision support service.
An existing database contains maximum doses for all drugs and recommended doses for patients with impaired renal function. In the decision support service ingredients are once again expressed as SNOMED CT substances.
The decision support platform will incorporate an alert filtering service in which physicians can set up their personal preferences for the displaying of alerts. For example, the dose form hierarchy of SNOMED CT will be used to enable filtering of unwanted alerts for specific dose forms (such as cutaneous dose forms).
1

A major application for Natural Language Processing technologies is indexing collections of free text transcripts or documents such that topic specific searches may be run on them. The challenge is to return ranked matches which permit selection of texts with high sensitivity and high specificity (i.e. that relevant documents are rarely overlooked and that irrelevant documents are rarely returned).
Clinical searches may be performed over transcripts or documents that reside in an electronic library, within medical records, or the Internet. Examples of searches include:
"Show me articles on this website concerned with inflammatory bowel disease"
"Does this patient have transcripts in their record suggesting a heart rhythm disturbance?"

Bevan Koopman's PhD thesis explores semantic and statistical approaches to search. The intention is to move beyond the limitations of plain keyword searching strategies for medical document retrieval. Characterizing these limitations as the 'semantic gap' Bevan identifies and addresses several issues including:
Vocabulary mismatch: hypertension vs. high blood pressure
Granularity mismatch: antipsychotic vs. haloperidol
Conceptual implication: e.g. from hemodialysis infer kidney failure
Inferences of similarity e.g. comorbidities (anxiety and depression)
His specific aim was to determine whether graph-based features and the propagation of information over a graph can provide an inference mechanism to bridge this semantic gap. As part of this work, he assessed the contribution of using SNOMED CT data within the graphs used to drive inferences.
The specific application in the thesis to find patients who match certain inclusion criteria for recruitment into clinical trials based on the analysis of free text transcripts from clinical records.
Queries included
Patients with depression on antidepressant medication
Patients treated for lower extremity chronic wound
Patients with AIDS who develop pancytopenia
Indexing methods were applied to the TREC MedTrack corpus - a standard collection of electronic texts containing de-identified reports from multiple hospitals in the United States. It includes nine types of transcripts: history and physical examinations, consultations, reports, progress notes, discharge summaries and emergency department operation reports, radiology, surgical pathology and cardiology reports. The collection as used contained around 100,000 reports within around 17,000 unique 'visits'.
Graphs have a number of characteristics that align with the requirements of semantic search as inference. The edges in a graph capture interdependence between concepts – which is identified as one of the semantic gap problems. Graphs are a common feature of both ontologies and retrieval models. The propagation of information over a graph — such as the popular PageRank algorithm used in Internet Search engines— provides a powerful means of identifying relevant information items (be they terms, concepts or documents). Ontologies such as SNOMED CT may also be represented as graphs.
The Graph Inference model developed by Bevan Koopman specifically addresses a number of semantic gap problems. Regarding vocabulary mismatch, the Graph Inference model utilizes a concept-based representation as this helps to overcome vocabulary mismatches (i.e. missed synonymy). The Graph Inference model specifically addresses granularity mismatch by traversing parent-child (i.e. 'is a') relationships.
The semantic gap problem of 'conceptual implication' is where the presence of certain terms in the document infer the query terms. For example, an organism may imply the presence of a certain disease. Such associations are captured in SNOMED CT and thus the Graph Inference model can specifically address conceptual implication by traversing those relationships.
Finally, the semantic gap problem of 'inference of similarity', where the strength of association between two entities is critical, is specifically addressed by the diffusion factor, which assigns a measure of similarity to each domain knowledge-based relationship. In the case of SNOMED CT the diffusion factor was derived from SNOMED CT relationships. It was noted that some relationships contributed to search sensitivity or conversely could lead to noise (loss of specificity) for the purpose of document retrieval. A weighting was applied (empirically) to each SNOMED CT relationship type and used as part of the relationship type component of the diffusion factor. For example, relationship type weightings included:
|is a| = 1.0
|active ingredient| = 1.0
|definitional manifestation| = 0.8
|associated finding| = 0.6
|severity| = 0.2
|laterality| = 0.2
Documents were parsed and analyzed using Lemur – a highly versatile and customizable open source information retrieval package developed at the University of Massachusetts. The construction of the graph was done using the open source LEMON graph library. The graph was serialized using LEMON and stored inside the Lemur index directory. For the MedTrack corpus, which was found to have a vocabulary size of 36,467 SNOMED CT concepts, the resulting graph was 4.4MB.
The findings of the thesis demonstrated that the graph based retrieval approaches using SNOMED CT derived data performed better than other approaches on 'hard queries'. A number of additional insights were also revealed. First, hard queries require inference and easy queries do not. Hard queries tended to be verbose and often contained multiple dependent aspects to the query (for example, a procedure and a diagnosis concept). Re-ranking using the Graph Inference model was effective here. Easy queries tended to have a small number of relevant documents and an unambiguous query concept. For these queries, inference was not required and the Bag-of concepts model was most effective. Overall, when valuable domain knowledge was provided by SNOMED CT, the Graph Inference model was effective — either by returning new relevant documents or by effectively re-ranking those selected. This again highlights the dependence on the underlying domain knowledge.
Regarding residual lack of sensitivity of all the IR strategies, Koopman suggests that an ideal ontology for information retrieval would not only contain definitional but also assertional data – for example "captopril can be used as a treatment of hypertension", "myocardial infarction [may] cause heart block" and "diabetes mellitus may lead to renal failure".
1