DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
| Section | ||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
These
instructions are for end users. With these instructions you can install
cTAKES, configure it, and use it to process text (typically text
associated with a medical record). If you were planning to expand,
change, or modify the code within cTAKES, refer to the cTAKES 2.5 6 Developer Install InstructionsGuide.
These instructions will cover installation and a test of the main product including trained models for sentence detection and tagging parts of speech, dictionaries from a subset of the UMLS, a very small subset of the full LVG resource, etc. Optional components will also be described.
...
Step | Example | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
1. Navigate to the source downloads for a released version on SourceForge | |||||||||||
2. Download the cTAKES-2.56.zip file. |
| ||||||||||
3.
Linux:
This folder we will call <cTAKES_HOME>. You will need to refer to the directory later. |
|
Process documents using cTAKES
...
Step | Example | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1. Open a command prompt and change to the cTAKES_HOME directory.
Linux:
|
| |||||||||||||||
2. Start the CAS Visual Debugger by running this command:
Linux:
The application may take a minute to start on slower hardware. |
| |||||||||||||||
3. An analysis engine (AE) needs to be loaded in order to process text.
Click Open. |
| |||||||||||||||
4. Copy the text in the example at the right (next cell) and paste the
| Dr. Nutritious | |||||||||||||||
3. From the menu bar, click Run -> Run AggregatePlaintextProcessor. |
| |||||||||||||||
4. Named entities are now recognized in this clinical document. |
|
...
Step | Example | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1. Open a command prompt and change to the cTAKES_HOME directory:
Linux:
|
|
...
2. Start the collection processing engine by running this command:
Linux:
The application may take a minute to start on slower hardware. |
| |||||||||||||||
3. This will bring up the Collection Processing Engine Configurator. In the Menu bar click File >Open CPE Descriptor |
| |||||||||||||||
4. Navigate to the file
Click Open. |
| |||||||||||||||
5. Click the Play button (green/blue play arrow near the bottom). |
| |||||||||||||||
6. You should see that one document was processed. You did process a collection of documents. In this case the collection only contained one just to show how to do it. Close the results window. |
| |||||||||||||||
7. Close the CPE application. You may be prompted to save changes. Since this was just a test you may click the No button. |
| |||||||||||||||
8. Open a new command prompt and change to the <cTAKES_HOME> | No example. | |||||||||||||||
9. To test the results there is a comparison tool that will help show that the results match expectations with the following syntax:
Where: <First File> is the first file to compare; <Second File> is the second file to compare; <diff-html> is where the results are written to | Windows:
Linux:
| |||||||||||||||
10. The resulting file will open for you. Look at the comparison to see the annotations resulting from this pipeline.
Linux:
|
|
...
Annotator | Description | Abbreviated | Example Analysis Engine (AE) | Example Collection processing Engine (CPE) | Example test data |
|---|---|---|---|---|---|
Clinical Document Pipeline | the complete cTAKES pipeline to obtain majority of cTAKES annotations | cdp | cTAKES_HOME/cTAKESdesc/cdpdesc/analysis_engine/AggregatePlaintextProcessor.xml | cTAKES_HOME/cTAKESdesc/cdpdesc/collection_processing_engine/test_plaintext.xml | cTAKES_HOME/testdata/cdptest |
Chunker | obtain cTAKES chunking annotations | chunker | cTAKES_HOME/cTAKESdesc/chunkerdesc/analysis_engine/ChunkerAggregate.xml | cTAKES_HOME/cTAKESdesc/chunkerdesc/collection_processing_engine/ChunkerCPE.xml | cTAKES_HOME/testdata/chunkertest |
Dependency Parser | obtain dependency parsing tree | dp | cTAKES_HOME/cTAKESdesc/dpdesc/analysis_engine/ClearParserTokenizedInfPosAggregate.xml | cTAKES_HOME/cTAKESdesc/dpdesc/collection_processing_engine/ClearParserCPE.xml | cTAKES_HOME/testdata/dptest |
Drug NER | the annotator to obtain drug annotations | drugner | cTAKES_HOME/cTAKESdesc/drugnerdesc/analysis_engine/DrugAggregatePlaintextProcesor.xml | cTAKES_HOME/cTAKESdesc/drugnerdesc/collection_processing_engine/DrugNER_PlainText_CPE.xml | cTAKES_HOME/testdata/drugnertest |
Dictionary Lookup | mapping cTAKES annotations to dictionaries (e.g., SNOMED_CT or RxNorm | lookup | cTAKES_HOME/cTAKESdesc/lookupdesc/analysis_engine/TestAggregateTAE.xml | cTAKES_HOME/cTAKESdesc/lookupdesc/collection_processing_engine/LookupCPE.xml | cTAKES_HOME/testdata/lookuptest |
PAD Term Spotter | identifying terms related to PAD | pad | cTAKES_HOME/cTAKESdesc/paddesc/analysis_engine/Radiology_TermSpotterAnnotatorTAE.xml | cTAKES_HOME/cTAKESdesc/paddesc/collection_processing_engine/Radiology_Sample.xml | cTAKES_HOME/testdata/padtest |
Smoking Status | the annotator to obtain document or patient-level smoking status | smoking | cTAKES_HOME/cTAKESdesc/smokingdesc/analysis_engine/SimulatedProdSmokingTAE.xml | cTAKES_HOME/cTAKESdesc/smokingdesc/collection_processing_engine/Sample_SmokingStatus_output_flatfile.xml | cTAKES_HOME/testdata/smokingtest |
Side Effect | the annotator to find side effect mentions and sentences from clinical documents | sideeffect | cTAKES_HOME/cTAKESdesc/sideeffectdesc/analysis_engine/SideEffectAggregateTAE.xml | cTAKES_HOME/cTAKESdesc/sideeffectdesc/collection_processing_engine/SideEffectCPE.xml | cTAKES_HOME/testdata/sideeffecttest |
Next Steps
The cTAKES 2.5 6 Component Use Guide will help you to understand in great detail each of the cTAKES components that have been installed. In some cases you can learn how to improve the components. However, before you go on to process text in production you will need to consider dictionaries and models.
...
Some models included in cTAKES may not represent your data distribution well. If you want to build or train your own models, please read the cTAKES 2.5 6 Component Use Guide, particularly:
- Training a sentence detector model
- Training a Part of Speech (POS) tagger model (Building a model Obtaining training data)
- Creating a Part of Speech (POS) tag dictionary (Building a tag dictionary)
- Training a chunker model (Building a model - Prepare GENIA training data)
- Training a dependency parser (Dependency Parser)