Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

For now, see: https://downloads.apache.org/tika/2.0.0/CHANGES-2.0.0.txt


Metadata

Removed duplicate/triplicate keys

Background: In early 1.x, we had basic metadata keys that were created somewhat ad hoc.  We then added normalized metadata keys based on standards such as Dublin Core, or we at least tried to add namespaces to the metadata keys for specific file formats.  To maintain backwards compatibility, we kept the old keys and added new keys.  This led to quite a bit of metadata bloat, where we'd have the same information two or three times.  In Tika 2.x, we slimmed down the metadata keys and relied only on, say Dublin Core if it exists.


Tika 1.xTika 2.x
Author, meta:author, dc:creatordc:creator
Last-Author, meta:last-authormeta:last-author
Creation-Date, date, dcterms:createddcterms:created
Last-Modified, modified, dcterms:modifieddcterms:modified
Last-Save-Date, meta:save-datemeta:save-date
Application-Name, extended-properties:Applicationextended-properties:Application
Character Count, meta:character-countmeta:character-count
Company, extended-properties:Companyextended-properties:Company
Edit-Time, extended-properties:TotalTimeextended-properties:TotalTime
Keywords, meta:keyword, dc:subjectmeta:keyword, dc:subject
Page-Count, meta:page-countmeta:page-count
Revision-Number, cp:revisioncp:revision
subject, cp:subject, dc:subjectdc:subject
Template, extended-properties:Templateextended-properties:Tempate
Word-Count, meta:word-countmeta:word-count


Changed Metadata Keys

Tika 1.xTika 2.x
X-Parsed-ByX-TIKA:Parsed-By




  • Metadata.RESOURCE_NAME_KEY has been renamed TikaCoreProperties.RESOURCE_NAME_KEY.
  • TikaCoreProperties.KEYWORDS has been renamed Office.KEYWORDS.
  • Meta X-Parsed-By has changed to X-TIKA:Parsed-By
  • X-TIKA:EXCEPTION:runtime has been changed to X-TIKA:EXCEPTION:container_exception

...