lunes, 11 de mayo de 2020

Generating FHIR Logical Models from archetypes: Transform and roll out!

Transforming archetypes to specific resources (like Observation or Questionnaire) can be useful for quick integrations. However, most of the remaining Resources are not as generic, and thus a specific transformation will not always be reusable enough to justify an specific transformation process. For this case, one thing we can do is to generate FHIR Logical Models. These are not intented to represent new Resources, but could be used to represent any archetype or template in a way that FHIR tooling can understand.

Generating Logical Models from archetypes or templates is a direct process, but in the end generation process differs a little from the already defined transformations to Observation or Questionnaire.

One of the main differences is that while the Observation transformation is based in generating the 'differential' part of the StructureDefinition, for Logical Models we actually generate the 'snapshot' part of the StructureDefinition.1

Another interesting thing is that this transformation process works regardless of the input archetype reference model. This is because process in completely based on the Archetype Object Model (AOM). 2

What this means is that this process can generate Logical Models not only from openEHR EHR Reference Model, but also from openEHR demographic model, ISO13606, HL7 CDA, or even CDISC ODM. In principle, generating archetypes works for any standard or local structures that are expressed as either XML Schemas or BMM and has the notion of "building block". The only real development needed to support a new reference model is to define a set of data types equivalences.

It is also worth noticing that this will work also for any kind of artifact that LinkEHR can import (OPT, OET, and even ADL2 archetypes), as they are all transformed to AOM structures in the end.

Process is quite straightforward:
  • As similar to other StructureDefinition generation, first the metadata part of the Logical Model is generated from the archetype/template.
  • Nodes in the archetype are traversed in order to generate the corresponding id, path, short text, definition, and min/max occurrences. Other attributes such as mustSupport, isModifier, and isSummary are also added.
  • Archetype internal references are transformed to contentReference
  • For types of Element and above, their type is transformed to BackboneElement.
  • Objects of type Element and their corresponding data type are fused as a single entity with the text, description, etc. of the Element and type of the transformed data type. 
  • Data type alternatives are added to the element path to create unique paths. Probably an invariant would be needed to show that they are actually alternatives.
There is still a few useful things missing in this transformation that could be added:
  • Probably mappings to the actual archetype paths can be defined, like it's already included in the Observation profiles. Currently they could be reverse engineered from the generated FHIR paths, but there is no reason to not include them as mappings (or even bindings).
  • Multilinguality is lost in this transformation. This probably differs quite a bit between STU3 and R4. Needs more examples on how it is usually done in FHIR.
  • Add the option to complete with reference model: Archetypes also define some kind of 'differential' view over the reference model, but it's easy to complete the archetype with the underlying reference model to get all constraints.
  • ArchetypeSlots are a completely different approach to using References to profiles. As far as I know, References can only be defined to point to specific profiles, while archetypes typical use case includes pointing to multiple ones (e.g. a given archetype or any of its specializations). This means that there is no direct mechanism to transform ArchetypeSlots for the moment.
All this process was already available on last week's LinkEHR release. Support to more data types has been added in this week's release. All feedback is welcome.
1. This wasn't always the case, as the original intent for Observation transformation was to let users choose if they wanted to define the differential or the snapshot+differential. This second option was discarded because it was really time consuming and FHIR tools can already generate the snapshot+differential from the differential
2. For FHIR Observation transformation the input could also be a ISO13606 archetype, but this was possible because the process looked for specific openEHR or ISO13606 classes.

lunes, 27 de abril de 2020

Generating FHIR Questionnaire resources

These days of quarantine (we are literally more than 40 days now here in Spain) have not been the most productive days of all, but I wanted to put a little effort in helping in the great effort that is the COVID openEHR template. Seriously, take a lot at it, it's impressive the quality and how fast deployment in real systems has been.

In any case, as I don't have a clinical background, I couldn't help much other than translating some not-so-clinical archetypes (occupation record, etc.) to Spanish. But one thing I could do was to finally work on another idea of generating Questionnaire instances from a given template/archetype. For that, openEHR COVID template seemed like the perfect use case: It would require fast iterations and no (human) time should be spent in the generation of the equivalent FHIR Questionnaire for each revision.

With that in mind, I looked at the available FHIR Questionnaire examples and took the common meta attributes and tried to relate them to the archetype ones. Some are always put as constants in any case (e.g. even if it was a published archetype, the autogenerated questionnaire always should be a 'draft').

After that two main big tasks were detected:

  1. Transform a given object in the archetype to the corresponding item in the questionnaire instance.
  2. Flatten the archetype hierarchy as much as possible to obtain all possible questions

For 1) the process was more or less direct: Generate a linkid from the template/archetype path, get the text from archetype, get the object occurrences to give value to the repeat attribute if necessary and finally transform the archetype rmtype to the corresponding questionnaire item type. For this, it was assumed that any remaining hierarchy objects should be treated as questionnaire "group". Data types are translated to the closest type. If an implicit value set is defined in a coded text of the archetype, the corresponding FHIR ValueSet is generated and the item points to it. In this case the type "choice" is selected.

For 2) an specific traverse algorithm for openEHR Reference model was designed. This algorithm traverses the template tree and only generates the items that are interesting (e.g. it avoids generating the History, Item_tree, etc. levels). This process also merges both Element and data type objects, so the same FHIR question item can have a question text and a data type.


Snippet of openEHR COVID-19 Pneumonia Diagnosis and Treatment (7th edition) template transformed into a FHIR Questionnaire



Besides these tasks, there were a few little challenges:
  • Even if you usually create profiles for each Resource, in case of questionnaires you define an instance of Questionnaire resource. Then your QuestionnaireResponse instances relate to that instance of the Questionnaire resource. This does not make the process more difficult, but means that not all the already developed methods to deal with the transformation of an archetype to a FHIR Observation Resource could be reused directly (e.g. different data types being allowed as parts of Observation and Questionnaire, which has already been discussed by Thomas Beale in the past).
  • Another thing that was available for the FHIR Observation transformation that is not available for Questionnaires is the mapping part. It should be very interesting to point to the exact archetype paths where a question comes from, but as far as I know this is not possible in vanilla questionnaires.
  • As already happened with the FHIR Observation transformation, FHIR doesn't really support the alternatives in monovaluated attributes. This is also true for FHIR Questionnaries and is also present in the COVID template (e.g. having an data value alternative of DV_TEXT and DV_CODED_TEXT, or DV_DURATION and DV_CODED_TEXT). As far as I see it there isn't a satisfactory solution to this yet, so for the moment only one of the alternatives is generated.
  • Archetype slots provide a challenge as they potentially reference an open set of archetypes, which as far as I understand the FHIR Reference Resource cannot handle. For the time being, ArchetypeSlots will be ignored in the transformation.
  • In case of openEHR COVID template, having the Symptom/Sign name as another element complicated things a little, as this should end merged with the corresponding values or presence absence. Luckily it was a pattern that could be easily implemented in step 2.
  • Archetypes are multilingual, but Questionnaires are not (maybe there is one extension out there that deals with it). In any case, getting the texts from a given language in the archetype vocabulary should be trivial, so multilingual questionnaires could easily become a reality with very little effort.
This functionality is part of LinkEHR since 2020-04-22. Try it and suggest improvements if you want :)

Next blog post will talk about the other new functionality introduced in that LinkEHR version, the generation of FHIR Logical Models from archetypes

jueves, 28 de noviembre de 2019

Extracting (all) CKM ValueSets

A derived results from the latest developments on the openEHR -> FHIR transformation is that I end up developing a method to analyze all the contents of an archetype to identify potential elements to end in a profile, regardless of the source class of the archetype. With that method developed, there was only a sensible thing to do...


Based on the CKM github mirror I applied the process to generate the profiles for each archetype, which in the end generates the ValueSets for each DV_CODED_TEXT, DV_ORDINAL, and even DV_QUANTITY units.

So this is the result. A set of 1023 ValueSet profiles where generated. This shows how much knowledge and effort has been put into developing CKM archetypes. For example, "location of measurement" ValueSet provides a set of 15 terms translated to 15 languages, which is something I've never seen before in a given subset.

Extract from locationofmeasurement ValueSet




I plan on regenerating this from time to time, so check the date for getting the latest version.

(and as a remainder, unless stated otherwise everything I put in this website has a CC-BY-SA license, so take it, try to break it and tell me if it actually breaks ;)

miércoles, 30 de octubre de 2019

Using LinkEHR to generate FHIR Observation profiles

After my last 'summer challenge' post one of the missing things was to make it publicly available in a simple way. This post shows how that algorithm was actually implemented in LinkEHR. This functionality is available in LinkEHR version that can be downloaded from linkehr.com

The process is quite straightforward:
  1. Open a given archetype (or import OPT)
    1. Process is NOT limited to openEHR Observation archetypes, any archetype can be feed as an input. Output will always be FHIR observation profile/s
  2. Launch the wizard in Advanced utilities -> Transform openEHR to FHIR observation
  3. Select an output path
  4. Select the archetype paths to be exported
    1. The process lists all paths with a leaf node, regardless of where are they included in the archetype (data, protocol, etc.)
  5. Select the type of the element in the FHIR profile (as Observation.component, as Observation.value or as a new extension)
  6. Check the suitable options from the bottom and press Finish



Process should provide a observation profile + a set of ValueSets coming from the coded values of the archetype. It also provides the Extension profiles from nodes selected as extension (if any)


What is currently supported via UI

  • Add translations to ValueSets
  • Support to OPTs as input 
    • Presented OPT paths are not as readable as in normal archetypes, but the process should work Ok anyway
  • Generate mapping file
    • Creates a LinkEHR mapping file containing the mapping the openEHR source paths to the target FHIR XML paths
  • Generate empty FHIR archetype
    • This creates an empty FHIR archetype for being use with the above mapping file. This allows the automatic generation of an XQuery program to transform openEHR data instances into FHIR data instances

 What is currently supported by code (hopefully soon in the UI)

  • Generate a set of observation profiles from a single openEHR archetype. 
    • Still evaluating how to create a simple UI for this functionality
  • Deal with underlying openEHR Reference Model
    • Some RM attributes are probably interesting to be included in profiles, which means some kind of merge of RM + archetype must be done
  • Generating STU3 profiles
    • It is supported to generate both R4 and STU3 "flavours"of FHIR. However, needs more testing to ensure everything is still correct after latest changes.

Currently working on

  • Dealing with DV_IDENTIFIER. 
    • No support for Identifier type in Observation.component.value[x] means that either identifiers are put in some specific paths in the FHIR profile or we create extensions. Trying to identify all these possible paths to ask users what they want to do with them

Potential improvements

  • Support other reference models
    • ISO 13606 would be almost trivial, but probably other RM such as HL7 CDA are interesting to tackle
  • Create Bundles to group the observations
  • Evaluate other classes (openEHR Composition with FHIR Composition resource)

Please tell me any doubt, problem, or error you found with the process so I can fix it. Probably best channel is by twitter (@yampeku)

PS: This functionality has actually been a month included in the tool, but only a few selected people knew it was already there to get their feedback

miércoles, 10 de julio de 2019

openEHR Observations to FHIR profiles, my 2019 summer challenge

This personal challenge started with a tweet from @siljelb asking for tooling for transforming openEHR archetypes to FHIR profiles. This was something I've always wanted to prove, so why not turn this into a personal summer challenge?

FHIR to openEHR 

 

My first thought was to generate an archetype from all the different components and extensions included in a FHIR profile. The rationale being that we already have systems capable of supporting arbitrary archetypes and archetypes have little problems when you want to include an unknown/big set of data values. After a little bit of profile analysis and a little trial and error I ended with a method that received an observation profile (with all the definitions of the extension it contained) and outputted a single observation archetype. As an example, here is the algorithm applied to the genomics profile.


Original genomics profile

Output observation archetype



(I would found later on that this particular example is not really accurate of how complex observations should be modeled in FHIR, but algorithm should stay more or less the same).

Although it seemed like a good idea at first, this transformation has a big issue: Two slightly different profiles will give us two different archetypes that will have to rely fully on terminology bindings to know they are talking about the same, which kind of goes against the philosophy of one archetype per concept. Archetypes are supposed to represent a universal use case (aiming to a max data set instead of a 80/20). These FHIR to openEHR transformations could work well enough in isolation or a single project but not a CKM-type environment. FHIR embraces local profiles and extensions, which could make their governance quite an issue.

openEHR to FHIR


One important lesson to get from this first experience is that a set of profiles will transform into a single archetype, but also a single archetype will end as a set of profiles! So while the FHIR → openEHR way only generates a single archetype and doesn't need any further user input besides choosing the set of profiles, the openEHR → FHIR transformation needs more information. Specifically, we need to tell the algorithm three things:
  • Select the Element that would go into the value part of the Observation profile (if any)
  • Select the list of elements that will end as components (they can go without value, they can also be an empty list)
  • Select which Elements of the original archetype are suitable to be transformed into extensions.
For first two, openEHR Clusters can give an idea of pieces of an observation that make sense as a group. So from there users should be able to choose what constitutes the value and what parts constitute the components. Clusters in Clusters are an interesting use case that probably needs that users choose directly how it should be generated. If we consider that every part in the archetype is equally important, then everything could end in profiles with values constraint to [0..0] and every archetype part modeled as components, which could make this process completely automatic.

For knowing what could be an extension, archetype slots provide an excellent indication. We could also analyze all archetypes in one go and create the different parts that could be reused as extensions (Can we create some default Protocol-like profiles? Do we need to take into account default parts from RM and always include them as extensions?)

Once we decide that, the generation of the Observation, Extension, and even ValueSet profiles is quite straightforward. For data types equivalences, I built this transformation based on the openEHR wiki page about data types equivalence. For the moment only the most usual data types have been translated (namely DV_TEXT, DV_CODED_TEXT, DV_DATE_TIME and DV_QUANTITY), but adding new ones should be easy (with maybe the exception of DV_ORDINAL alternatives).

Snippet of State of Dress valueset


The profile differences between STU3 and R4 aren't too many, so I've included support for both generations. User can choose which one generate as a parameter.

The generated profiles were adjusted thanks to the Hammer tool, so they should be correct for STU3. After a bit of tweaking I got rid of the unknown error that was preventing me to import them into Forge R4, so they should also be R4 compliant.

Weight profile from openEHR weight observation archetype


What's next

  • More testing! 
    • Do you have any archetype you want to test?
    • Could be interesting to process all CKM in one go? (at least for ValueSets)
  • Should we put all generated ValueSets in our FHIR terminology service?
  • More data types!
    • Some data types are still missing. I think that it is possible that some data types, such as openEHR Ordinals alternatives will provide bigger challenge if they have additional optional elements.
  • Make it more configurable, e.g. change base uris, etc.
  • Put the algorithm into LinkEHR Editor and provide a minimal user interface to ease the element selection process.
  • Use the process to generate the mapping that will generate automatically a transformation program to generate the different FHIR Observation instances.
    • Mapping seems that could be bidirectional, as it seems that all mappings are atomic.

Lessons learned

  • It is feasible to create a semiautomatic translation from openEHR to FHIR. We can make global assumptions to make this process automatic.
  • FHIR to openEHR transformation is straightforward, however to take advantage of the big amount of high quality clinical models already avaliable in openEHR it seems preferable to use openEHR archetypes as a basis and generate profiles for them.
  • The differences in StructureDefinition between STU3 and R4 that matter to this transformation are manageable.

Bonus for those of you getting to the end, the good ol' Blood pressure archetype as a FHIR observation profile

miércoles, 9 de mayo de 2018

Using SNQuery to test FHIR subsets

When defining Snomed CT subsets, the most common approach is to define them in an extensional way. This is also the case for national bodies or subsets defined in FHIR specifications. These extensional subsets definition have several potential problems, such as missing out concepts out of the subset or the existence of homograph words that cause the choosing of wrong terms. There is a missing opportunity for the use of Snomed Expressions Constraints for subset definitions. We will use SNQuery to demonstrate how we can create or validate existing subsets in an easy way.

Subsets with logical definition (is-a)


The first example, is those subsets which also contain a logical definition with "is-a" relationships such as medication codes valueset or body site valueset
For medication codes valueset, the definition is as follows:

This definition can be easily translated to Expression Constraint Language like this
<< 410942007 |Drug or medicament (substance)| OR << 373873005 |Pharmaceutical / biologic product (product)| OR << 106181007 |Immunologic substance (substance)|
 This expression in the latest substrate available (20180131) contains 28928 concepts

Lists of codes


Several FHIR subsets are defined as sets of codes (such as bodysite relative location). The main problem with this approach is that subsets may be incomplete, or be potentially wrong, as hand picked words could have several meanings (e.g. the same term can be used to represent the procedure and the tissue where the procedure is made). These errors are revealed easily by looking to some graphs. In case of the bodysite relative location, the equivalent expression is the following one:
419161000 or 419465000 or 51440002 or 261183002 or 261122009 or 255561001 or 49370004 or 264217000 or 261089000 or 255551008 or 351726001 or 352730000






This graph shows where most of the concepts fall. In this case, all bodysite relative locations are qualifier values, which seems correct.

These kinds of visualizations become more useful, the more concepts the subset has. For example facility codes subset, which contains 79 concepts.





The last graph can be interpreted as follows: Of the 79 concepts contained in the subset  94% (74 out of 79) are a 'site of care', with 4 of the remaining ones being 'community environment' and a single one is a 'hospital environment'. With this kind of visualization, some questions can be raised: That single code in the hospital environment subtree should be referring to itself and all the allowed children? Could we get away with simplifying the subset to the expression "< 276339004 |Environment (environment)|" which contains all the environments known in Snomed? Do the terms annotated in the original FHIR subset with "--OTHER--NOT LISTED" should always be translated into a children or self operation?

Validating subsets

One of the advantages of this approach is that the graphical representation also allows for a quick review of the quality of the proposed subset. As an example the specimen collection method, which can be defined with the expression constraint

119295008 |Specimen obtained by aspiration| OR 413651001 |Bioptics| OR 360020006 |Extirpation - action| OR 430823004 |Examination of midstream urine specimen| OR 16404004 |Induced| OR 67889009 |Irrigation| OR 29240004 |Autopsy examination| OR 45710003 |Sputum| OR 7800008 |Punctate| OR 258431006 |Scrapings| OR 20255002 |Blushing| OR 386147002 |Smear procedure| OR 278450005 |Finger stick|

This subset contains 13 concepts, and shows the following graph:




One thing that can be quickly seen is that the focus concept (which can be interpreted as the minimum common ancestor) is the Snomed CT root concept. This means that concepts in the subset have no other common parent aside from the Snomed root concept. This serves as a sign that subset definition has potential problems.

A more detailed analysis shows that:
  • 5 / 12 concepts are in the procedure hierarchy, which seems fitting as we are talking about collection methods.
  • 3 / 12 concepts are in the qualifier value hierarchy
  • 2/ 12 concepts are in the specimen hierarchy
  • 1 / 12 concept is in the substance hierarchy
  • 1 / 12 concept is in the observable entity hierarchy 
  • 1 concept is inactive and shouldn't be used

 Taking a careful look at the concepts not in procedure hierarchy, some problems become apparent:
  • The concept "45710003 |Sputum (substance)|", which refers to the substance, is used to refer to the method of collection. In this case, the concept "37705003 |Collection of sputum (procedure)|" seems way more fitting to the purpose of the subset
  • The concept "20255002 |Blushing, function (observable entity)|"refers to an observable entity. It seems that the correct concept could be "225063006 | Flushing cannula (procedure) |"
  • Similar to the last one, "258431006 |Scrappings (specimien)|" and "Specimen obtained by aspiration (specimen)" seem to be unfitting for the purpose of the subset. Probably, "56757003 |Scraping (procedure)|" and "14766002 | Aspiration (procedure) |" should be used instead. 
  • The inactive concept "386147002 | Smear procedure (procedure) |" was made inactive because it was ambiguous. By reviewing the code in the Snomed Browser (see RefSet tab) we can see that code should be "448895004 | Sampling for smear (procedure) |" or "448938001 | Preparation of smear (procedure) |" (or both).
  • Regarding the selected qualifier values, the correct approach shouldn't be to add these qualifiers to the set, but to allow the procedures that use this qualifier value. This can be expressed as 71388002|Procedure (procedure)| which 260686004 |Method (attribute)| is any of the three qualifier values. Formally, the expression for these qualifier values would end as:
 <71388002|Procedure (procedure)|:260686004 |Method (attribute)|=(16404004 |Induced (qualifier value)| OR 360020006 |Extirpation - action (qualifier value)| OR 7800008 |Punctate (qualifier value)|)
Note: Even if Induced and Punctuate terms are not currently used as destination of any attribute in Snomed CT, this expression constraint allows to validate the post-coordination of procedures that use these methods
In addition to validating the subset we should also ask ourselves again if there is some expression that could be used to group terms, and by doing so, simplify the expression.

Conclusion 

As the examples show, these subsets' visualizations could help clinical experts in the subset definition, validation, and curation. Extensional subsets may seem easy to define, but could led to potential problems if the hierarchy is not taken into account when defining the subset. Even if (arguably) one of the advantages of defining extensional subsets is to limit the possible inputs in a form, official provided subsets should always try to include all and every term useful for the subset to ease interoperability. When implementing the subset in a given organization is always better to further refine a subset than to extend it with terms not originally included in it.




martes, 23 de enero de 2018

Having fun with Snomed expression constraints (and learning something in the meantime)

This article wants to be a fun introduction to the Snomed expression constraint language in order to show its capabilities. This article assumes no prior knowledge of the expression constraint language, so it will start with a little introduction to it (if you already know about the Snomed expression constraint language should be safe to jump to point 2). In this post both IHTSDO Snomed browser and VeraTech SNQuery will be used.

Snomed Expression Constraint Language basics

In this section a few operators from the Snomed Expression Constraint Language will be explained. For a complete explanation of the Snomed Expression Constraint Language visit the official documentation.

Simple expression constraints

The following simple operators already provide great functionality for querying Snomed hierarchy:

Descendant of: The constraint is satisfied by all the transitive descendants of a given Snomed concept. This is denoted by the operator 'less than' (<). For example, the expression
< 64572001 | Disease (disorder) | provides about 74k concepts which includes concepts such as Anemia, Hematoma, or Inflammatory fibroid polyps of stomach

Descendant or self: Similar to 'descendants of' operator, the operator Descendant or self (denoted by two 'less than'  symbols) is satisfied by all the transitive descendants of a given Snomed concept plus the concept itself. For example, << 11466000 |Cesarean section (procedure)| includes both the descendants of cesarean section and the cesarean section term itself.

There are more simple operators, but probably these two are the most used by far.

Refinements

A refinement in a Snomed expression allows the filtering the resulting set using one or more attribute constraints.

One of the great things about Snomed is that terms themselves can be defined by refining existing terms (see Snomed compositional grammar). E.g. Hepatitis A (40468003) is a disease (64572001) found (363698007) at the liver structure (10200004) with a inflammation (23583003) morphology (116676008) caused by (246075003) Hepatitis A virus (32452004). Note that these expressions contain both clinical terms such as hepatitis A or disease, but also attributes such as associated morphology and finding site, which have their own Snomed codes. These attributes can be used to refine defined sets.

Attribute refinement restricts the meaning set of clinical meanings to those satisfying the refinement condition. Similarly to the Snomed compositional grammar, a 'colon' (':') is used in the expression.

As an example, pulmonar diseases could be defined as <64572001 |Disease (disorder)| : 363698007 |Finding site| = 39607008 |Lung structure (body structure)|

Note that these attributes have a "direction": the above expression returns all the diseases whose finding site is the lung structure. There will be times where we want to select the target term of a relationship and constraint the source. We can achieve this by using the 'reverse' operator ('R'). E.g with <64572001 |Disease (disorder)| : 246075003 |Causative agent| = 49872002 |Virus (organism)|  we can obtain all the diseases caused by viruses, and by reversing the attribute with an expression such as <49872002 |Virus (organism)| : R 246075003 |Causative agent| = <64572001 |Disease (disorder)| the subset of the viruses that cause diseases can be obtained.

There are more ways to refine an Snomed expression, but with these basic ones we can start 'playing' with Snomed.

Having Fun

With these operators and refinements in mind, we can start navigating Snomed without a deep knowledge of the underlying Snomed conceptual model, i.e. what attributes are valid in each hierarchy.

Finding what to look for

We defined an application that needed a list of cancer diagnosis (coded in ICD-10), but also a location where these cancer were found. Could we use Snomed to provide us with a (tentative) set of terms to fill this field?

Even if you don't know Snomed conceptual model, you probably know examples of what you are looking for. I will use 'lung cancer' as an example.

Navigating Snomed

Searching the term in the Snomed browser allows us to dig the different terms that make up the term meaning. 'Lung cancer' is a synonym of 'Primary malignant neoplasm of lung (disorder)', with term
9388000.

In the concept details we can examine the Expression' tab and then look at 'Expression from Stated Concept Definition'. That expression precisely defines the Snomed term from other existing Snomed terms. In this case we want to know which attributes are valid, in our case 'disorders'. By that definition, lung cancer (93880001) is a disease (64572001) found (363698007) at the lung structure (39607008) with a malignant neoplasm (86049000) morphology (116676008). We can generalize that expression by navigating Snomed hierarchy. For example, we could ask for all the primary cancers that have a finding site in any body structure <64572001 |Disease (disorder)| :{363698007 |Finding site| = <123037004 | Body structure (body structure) |,  116676008 |Associated morphology| = 86049000 |Malignant neoplasm, primary (morphologic abnormality)|} which results in a subset of ~600 terms.
An alternative is to look for all primary, secondary, or other cancers with a finding site in any body structure <363346000 |Malignant neoplastic disease (disorder)| :363698007 |Finding site| = <123037004 | Body structure (body structure) | which contains ~3700 terms.

In addition to give us this subset list, SNQuery allows us to simplify the expressions in order to reduce expression processing time. These simplified queries return the same terms (same subset) but contain more precise codes that ease the expression or makes it more precise and clearer. As an example, the expression <64572001 |Disease (disorder)| :{363698007 |Finding site| = <123037004 | Body structure (body structure) |,  116676008 |Associated morphology| = 86049000 |Malignant neoplasm, primary (morphologic abnormality)|} can be simplified as  <372087000 |Primary malignant neoplasm (disorder)|:{363698007 |Finding site| = <123037004 |Body structure (body structure)|, 116676008 |Associated morphology| = 86049000 |Malignant neoplasm, primary (morphologic abnormality)|}. First expression needs more than 5 seconds to be processed, while the second expression is about 150 milliseconds

Going in reverse

Once we have found a suitable expression we can reverse it to get the desired results. In our case, instead of looking for 'all primary, secondary, or other cancers with a finding site in any body structure' we will reverse it to express 'all body structures thar are finding sites in primary, secondary, or other cancers'. This subset can be expressed as <91723000 |Anatomical structure (body structure)|: R 363698007 |Finding site (attribute)| =<363346000 |Malignant neoplastic disease (disorder)| and contains ~880 terms

We could use this subset list as a first approach to populate our user interface and add Snomed codes into the mix. If we have a ICD code we could potentially use the official Snomed mapping and use it to validate the other fields (in this case, diagnosis with their location).