While users can run tika-eval-app on their own machines with their own documents, the Apache Tika, Apache PDFBox and Apache POI communities have gathered >1TB of documents from govdocs1 and from Common Crawl to serve as a regression testing corpus. Before a release, we'll run the last release against the candidate release to identify potential regressions.
This page is intended for committers/PMC members with access to the VM who want to run the regression tests. The example focuses on testing a SNAPSHOT version of PDFBox, but the steps are nearly identical for the full Tika eval or for sub projects. See TikaEval for more information on the tika-eval-app module by itself. See this blog for a description of running this project on Tika's VM.
The driver appBatchExecutor.sh, the various configuration files and the file lists for PDFs are available here: batch-scripts.tgz.
If you haven't done so in your .bashrc file, make sure to umask g+rw before running anything
The main working directory is: /data1/tools/tika/batch
jai-imageio-jpeg2000-1.4.0.jar, sqlite-jdbc-3.46.0.0.jar and zstd-jni-1.5.6-4.jar/data1/tools/tika/batch/logs/data1/tools/tika/batch/nohup.out/data1/tools/tika/batch/binappBatchExecutor.sh to-o /data1/extracts/pdfboxA-fileList fileLists/ccAndBugTracker_pdfs.txtnohup ./appBatchExecutor.sh &mvn clean installmvn clean on the whole Tika project and make sure that your IDE has picked up the changesmvn clean install/data1/tools/tika/batch/bin, rename the existing nohup.out to nohup-A.out, rename logs/ to logs-A//data1/tools/tika/batch/binappBatchExecutor.sh to-o /data1/extracts/pdfboxB-fileList fileLists/ccAndBugTracker_pdfs.txtnohup ./appBatchExecutor.sh &/data1/tools/tika/eval, remove the existing db file pdfboxAvsB.mv.db if you don't want to rename it.nohup java -jar tika-eval-app-X.Y.jar Compare -extractsA /data1/extracts/pdfboxA -extractsB /data1/extracts/pdfboxB -db pdfboxAvsB&reports/: rm -r reportsjava -Djava.io.tmpdir=tmp -jar tika-eval-app-X.Y.jar Report -db pdfboxAvsB – Note the -Djava.io.tmpdir=tmp – need to set the tmp directory to something writeable by 'collab'When this process completes, you'll have all of the reports written to /data1/tools/tika/eval/reports/.
With the expansion of the regression corpus, I'm finding that H2 isn't able to write the reports – no matter the -Xmx, even after a few hours, it doesn't even get to the point of creating the reports directory.
I should set up postgres on our VM, but I haven't gotten around to that yet. For now, I'm copying the H2 db to Postgresql and then writing the reports from there. The code to copy H2->postgres is available here: tika-addons.
I had to modify the report SQL slightly to work with Postgresql, and I stripped out some of the reports/calculations that aren't critical to the full regression tests. The modified report SQL is available comparison-reports-pg.xml