|
These instructions are under construction. |
These instructions are for end users. With these instructions you can install cTAKES, configure it, and use it to process text (typically text associated with a medical record). If you were planning to expand, change, or modify the code within cTAKES, refer to the cTAKES 3.0 Developer Guide.
These instructions will cover installation and a test of the main product including trained models for sentence detection and tagging parts of speech, dictionaries from a subset of the UMLS, a very small subset of the full LVG resource, etc. Optional components are described in the Component Use Guide.
Once you have finished installing cTAKES, you will be able to see what cTAKES is capable of. Further exploitation of the software's ability may require following a few additional steps involving what dictionaries are being used. The last section on this page covers these next steps.
Step |
Example |
|||
|---|---|---|---|---|
1. Make sure you have Java 1.6 or higher. Most systems come with Java already installed.
|
Windows:
|
Step |
Example |
|||
|---|---|---|---|---|
1. Navigate to the cTAKES downloads page on the Apache site and download the binary package. Select a mirror site and press the Change button to modify the URL to your desired mirror location before doing the download or accept the default.
|
Windows:
|
|||
2. (Optional but recommended) Verify the downloaded files against a file signature to ensure you have the proper and complete file. |
No example |
|||
3. Unzip the file you downloaded into a directory that you want to be the cTAKES install location. The compressed files contain a single directory at the top level. This folder we will call <cTAKES_HOME>. You will need to refer to this directory later.
Linux:
|
Windows:
|
|||
4. Download cTAKES 3.0 Dictionaries and models.
Go to http://sourceforge.net/projects/ctakesresources/files/ and download the ZIP file with a matching version from the ctakesresources project. |
Windows:
Linux:
|
|||
5. Copy (or move) the resources to cTAKES_HOME.
|
Windows:
Linux:
|
This version allows you to test most components bundled in cTAKES in two different ways:
You will need a windowing environment on Linux to run these tools.
Step |
Example |
|||
|---|---|---|---|---|
1. Open a command prompt and change to the cTAKES_HOME directory.
|
Linux:
|
|||
2. Start the CAS Visual Debugger by running this command: |
Windows:
Linux:
|
|||
3. Copy the example text from the next cell in this table and paste the contents into the Text section of CVD, replacing the text that is already there.
|
|
|||
4. An analysis engine (AE) needs to be loaded in order to process text.
|
|
|||
5. From the menu bar, click Run -> Run AggregatePlaintextProcessor. |
|
|||
6. You'll get a list of all the annotations for this clinical document in the Analysis Results frame. Annotations such as concepts mentioned, division by sentence, etc from the pipeline are viewable. To see one, in the Analysis Results frame, click on the key in front of:
This will show an AnnotationIndex in the lower frame. Select any annotation in that lower frame and you will see the text discovered in
Now select items in the lower frame to see the text being annotated. |
|
Step |
Example |
|||
|---|---|---|---|---|
1. Open a command prompt and change to the cTAKES_HOME directory:
|
Linux:
|
|||
2. Start the collection processing engine by running this command: |
Windows:
Linux:
|
|||
3. This will bring up the Collection Processing Engine Configurator. In the Menu bar click File >Open CPE Descriptor |
|
|||
4. Navigate to the file
Click Open.
|
No example |
|||
5. Click the Play button (green/blue play arrow near the bottom). |
|
|||
6. You should see that one document was processed. You did process a collection of documents. In this case the collection only contained one just to show how to do it. Close the results window. |
|
|||
7. Close the CPE application. You may be prompted to save changes. Since this was just a test you may click the No button. |
|
|||
8. Open a new command prompt and change to the <cTAKES_HOME> |
|
|||
9. To test the results there is a comparison tool that will help show that the results match expectations with the following syntax:
Where: <First File> is the first file to compare; <Second File> is the second file to compare; <diff-html> is where the results are written to |
Windows:
|
|||
10. The resulting file will open for you. Look at the comparison to see the annotations resulting from this pipeline.
Linux:
|
|
Using the same CVD and CPE programs in the manner described above, you can test all the other components. The analysis engines and collection processing engines shipped with cTAKES for some of the annotators are described in the following table.
Annotator |
Description |
Abbreviated |
Example Analysis Engine (AE) |
Example Collection processing Engine (CPE) |
Example test data |
|---|---|---|---|---|---|
Clinical Document Pipeline |
the complete cTAKES pipeline to obtain majority of cTAKES annotations |
cdp |
cTAKES_HOME/desc/ctakes-clinical-pipeline/desc/analysis_engine/AggregatePlaintextProcessor.xml |
cTAKES_HOME/desc/ctakes-clinical-pipeline/desc/collection_processing_engine/test1.xml |
cTAKES_HOME/testdata/cdptest |
Chunker |
obtain cTAKES chunking annotations |
chunker |
cTAKES_HOME/desc/ctakes-chunker/desc/ChunkerAggregate.xml |
cTAKES_HOME/desc/ctakes-chunker/desc/ChunkerCPE.xml |
cTAKES_HOME/testdata/chunkertest |
Dependency Parser |
obtain dependency parsing tree |
dp |
cTAKES_HOME/desc/ctakes-dependency-parser/desc/analysis_engine/ClearParserSRLTokenizedInfPosAggregate.xml |
cTAKES_HOME/desc/ctakes-dependency-parser/desc/collection_processing_engine/ClearParserTestCPE.xml |
cTAKES_HOME/testdata/dptest |
Drug NER |
the annotator to obtain drug annotations |
drugner |
cTAKES_HOME/desc/ctakes-drug-ner/desc/analysis_engine/DrugAggregatePlaintextProcesor.xml |
cTAKES_HOME/desc/ctakes-drug-ner/desc/collection_processing_engine/DrugNER_PlainText_CPE.xml |
cTAKES_HOME/testdata/drugnertest |
Dictionary Lookup |
mapping cTAKES annotations to dictionaries (e.g., SNOMED_CT or RxNorm |
lookup |
cTAKES_HOME/desc/ctakes-dictionary-lookup/desc/analysis_engine/TestAggregateTAE.xml |
cTAKES_HOME/desc/ctakes-dictionary-lookup/desc/collection_processing_engine/LookupCPE.xml |
cTAKES_HOME/testdata/lookuptest |
PAD Term Spotter |
identifying terms related to PAD |
pad |
cTAKES_HOME/desc/ctakes-pad-term-spotter/desc/analysis_engine/Radiology_TermSpotterAnnotatorTAE.xml |
cTAKES_HOME/desc/ctakes-pad-term-spotter/desc/collection_processing_engine/Radiology_Sample.xml |
cTAKES_HOME/testdata/padtest |
Smoking Status |
the annotator to obtain document or patient-level smoking status |
smoking |
cTAKES_HOME/desc/ctakes-smoking-status/desc/analysis_engine/SimulatedProdSmokingTAE.xml |
cTAKES_HOME/desc/ctakes-smoking-status/desc/collection_processing_engine/Sample_SmokingStatus_output_flatfile.xml |
cTAKES_HOME/testdata/smokingtest |
Side Effect |
the annotator to find side effect mentions and sentences from clinical documents |
sideeffect |
cTAKES_HOME/desc/ctakes-side-effect/desc/analysis_engine/SideEffectAggregateTAE.xml |
cTAKES_HOME/desc/ctakes-side-effect/desc/collection_processing_engine/SideEffectCPE.xml |
cTAKES_HOME/testdata/sideeffecttest |
The cTAKES 3.0 Component Use Guide will help you to understand, in great detail, each of the cTAKES components that have been installed. In some cases you can learn how to improve the components.
Also, before you go on to process text in production you will need to consider dictionaries and models. cTAKES does not distribute from Apache a complete dictionary capable of annotating production data. The models provided have been trained on data that may not match your data well enough to be effective. In most cases, you will need to modify the dictionaries and train models on your own data to be effective.