|
|
These instructions are for end users. With these instructions you can install cTAKES, configure it, and use it to process text (typically text associated with a medical record). If you were planning to expand, change, or modify the code within cTAKES, refer to the cTAKES 2.6 Developer Guide.
These instructions will cover installation and a test of the main product including trained models for sentence detection and tagging parts of speech, dictionaries from a subset of the UMLS, a very small subset of the full LVG resource, etc. Optional components will also be described.
Once you have finished installation of cTAKES, you will be able to see what cTAKES is capable of. Further exploitation of the software's ability may require following a few additional steps involving what dictionaries are being used. These are the last steps in these instructions.
Step |
Example |
||
|---|---|---|---|
1. Make sure you have Java 1.6 or higher. Most systems come with Java already installed.
If you do not you can install Java from java.com. |
|
Step |
Example |
||
|---|---|---|---|
1. Navigate to the source downloads for a released version on SourceForge |
|
||
2. Download the cTAKES-2.6.zip file. |
|
||
3.
Linux:
This folder we will call <cTAKES_HOME>. You will need to refer to the directory later. |
|
This version allows you to test most components bundled in cTAKES in two different ways:
Step |
Example |
|||
|---|---|---|---|---|
1. Open a command prompt and change to the cTAKES_HOME directory.
Linux:
|
|
|||
2. Start the CAS Visual Debugger by running this command:
Linux:
The application may take a minute to start on slower hardware. |
|
|||
3. An analysis engine (AE) needs to be loaded in order to process text.
Click Open. |
|
|||
4. Copy the text in the example at the right (next cell) and paste the
|
Dr. Nutritious |
|||
3. From the menu bar, click Run -> Run AggregatePlaintextProcessor. |
|
|||
4. Named entities are now recognized in this clinical document. |
|
Step |
Example |
|||
|---|---|---|---|---|
1. Open a command prompt and change to the cTAKES_HOME directory:
Linux:
|
|
|||
2. Start the collection processing engine by running this command:
Linux:
The application may take a minute to start on slower hardware. |
|
|||
3. This will bring up the Collection Processing Engine Configurator. In the Menu bar click File >Open CPE Descriptor |
|
|||
4. Navigate to the file
Click Open. |
|
|||
5. Click the Play button (green/blue play arrow near the bottom). |
|
|||
6. You should see that one document was processed. You did process a collection of documents. In this case the collection only contained one just to show how to do it. Close the results window. |
|
|||
7. Close the CPE application. You may be prompted to save changes. Since this was just a test you may click the No button. |
|
|||
8. Open a new command prompt and change to the <cTAKES_HOME> |
|
|||
9. To test the results there is a comparison tool that will help show that the results match expectations with the following syntax:
Where: <First File> is the first file to compare; <Second File> is the second file to compare; <diff-html> is where the results are written to |
Windows:
|
|||
10. The resulting file will open for you. Look at the comparison to see the annotations resulting from this pipeline.
Linux:
|
|
Using the same CVD and CPE programs in the manner described above, you can test all the other components. The analysis engines and collection processing engines shipped with cTAKES for some of the annotators are described in the following table.
Annotator |
Description |
Abbreviated |
Example Analysis Engine (AE) |
Example Collection processing Engine (CPE) |
Example test data |
|---|---|---|---|---|---|
Clinical Document Pipeline |
the complete cTAKES pipeline to obtain majority of cTAKES annotations |
cdp |
cTAKES_HOME/cTAKESdesc/cdpdesc/analysis_engine/AggregatePlaintextProcessor.xml |
cTAKES_HOME/cTAKESdesc/cdpdesc/collection_processing_engine/test_plaintext.xml |
cTAKES_HOME/testdata/cdptest |
Chunker |
obtain cTAKES chunking annotations |
chunker |
cTAKES_HOME/cTAKESdesc/chunkerdesc/analysis_engine/ChunkerAggregate.xml |
cTAKES_HOME/cTAKESdesc/chunkerdesc/collection_processing_engine/ChunkerCPE.xml |
cTAKES_HOME/testdata/chunkertest |
Dependency Parser |
obtain dependency parsing tree |
dp |
cTAKES_HOME/cTAKESdesc/dpdesc/analysis_engine/ClearParserTokenizedInfPosAggregate.xml |
cTAKES_HOME/cTAKESdesc/dpdesc/collection_processing_engine/ClearParserCPE.xml |
cTAKES_HOME/testdata/dptest |
Drug NER |
the annotator to obtain drug annotations |
drugner |
cTAKES_HOME/cTAKESdesc/drugnerdesc/analysis_engine/DrugAggregatePlaintextProcesor.xml |
cTAKES_HOME/cTAKESdesc/drugnerdesc/collection_processing_engine/DrugNER_PlainText_CPE.xml |
cTAKES_HOME/testdata/drugnertest |
Dictionary Lookup |
mapping cTAKES annotations to dictionaries (e.g., SNOMED_CT or RxNorm |
lookup |
cTAKES_HOME/cTAKESdesc/lookupdesc/analysis_engine/TestAggregateTAE.xml |
cTAKES_HOME/cTAKESdesc/lookupdesc/collection_processing_engine/LookupCPE.xml |
cTAKES_HOME/testdata/lookuptest |
PAD Term Spotter |
identifying terms related to PAD |
pad |
cTAKES_HOME/cTAKESdesc/paddesc/analysis_engine/Radiology_TermSpotterAnnotatorTAE.xml |
cTAKES_HOME/cTAKESdesc/paddesc/collection_processing_engine/Radiology_Sample.xml |
cTAKES_HOME/testdata/padtest |
Smoking Status |
the annotator to obtain document or patient-level smoking status |
smoking |
cTAKES_HOME/cTAKESdesc/smokingdesc/analysis_engine/SimulatedProdSmokingTAE.xml |
cTAKES_HOME/cTAKESdesc/smokingdesc/collection_processing_engine/Sample_SmokingStatus_output_flatfile.xml |
cTAKES_HOME/testdata/smokingtest |
Side Effect |
the annotator to find side effect mentions and sentences from clinical documents |
sideeffect |
cTAKES_HOME/cTAKESdesc/sideeffectdesc/analysis_engine/SideEffectAggregateTAE.xml |
cTAKES_HOME/cTAKESdesc/sideeffectdesc/collection_processing_engine/SideEffectCPE.xml |
cTAKES_HOME/testdata/sideeffecttest |
The cTAKES 2.6 Component Use Guide will help you to understand in great detail each of the cTAKES components that have been installed. In some cases you can learn how to improve the components. However, before you go on to process text in production you will need to consider dictionaries and models.
cTAKES includes the complete UMLS (SNOMED-CT and RxNorm) dictionaries.
To use them, you must have a UMLS username and password, and an Internet connection.
If you do not have a UMLS username and password, you may request one at UMLS Terminology Services |
In order to use the UMLS dictionaries shipped with cTAKES you will need to do two things:
(1) Change the UMLSUser and UMLSPW <nameValuePair> strings in these descriptor files with your UMLS username and password.
The following shows where in the files you would make the changes. (Do not change the <configurationParameters> by the same name.)
<nameValuePair> <name>UMLSUser</name> <value> <string>YOUR_UMLS_USERNAME_HERE</string> </value> </nameValuePair> <nameValuePair> <name>UMLSPW</name> <value> <string>YOUR_UMLS_PASSWORD_HERE</string> </value> </nameValuePair> |
(2) Include the DictionaryLookupAnnotatorUMLS.xml Analysis Engine within your aggregate Analysis Engine or switch to the ones provided by cTAKES. cTAKES has provided duplicates of shipped Analysis Engine descriptors, put UMLS in the name, and placed DictionaryLookupAnnotatorUMLS.xml within them for these components:
So you simply need to switch to using those descriptors. For example, if you were using AggregateCdaProcessor.xml in the Clinical Documents pipeline you would switch to using AggregateCdaUMLSProcessor.xml instead and you will now hook into the complete dictionaries.
You can, of course, modify your own aggregate Analysis Engine files and place the DictionaryLookupAnnotatorUMLS.xml Analysis Engine within them.
Since this is an in-memory database implementation, please be patient during the initial load as it could take approximately 20-30 seconds for the database to initialize.
If you would like to go back to using the small sample dictionaries that do not require a UMLS username, use the DictionaryLookupAnnotator.xml (UMLS is not in the file name) Analyis Engine descriptor in your aggregate. Just removing your password from the DictionaryLookupAnnotatorUMLS.xml files will not switch you back to the small sample dictionaries.
We have successfully tested the 2008 release of the full LVG data. In order to use this release of the full LVG data you should:
To install customized dictionaries for RxNorm, SNOMED-CT, or other vocabularies that are available through the UMLS, see the following posts on the cTAKES forums:
Some models included in cTAKES may not represent your data distribution well. If you want to build or train your own models, please read the cTAKES 2.6 Component Use Guide, particularly: