[jira] [Commented] (TIKA-2224) OneNote formats support - Mime Magic and Parser

classic Classic list List threaded Threaded
1 message Options
Reply | Threaded
Open this post in threaded view
|

[jira] [Commented] (TIKA-2224) OneNote formats support - Mime Magic and Parser

Markus Jelsma (Jira)

    [ https://issues.apache.org/jira/browse/TIKA-2224?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16994717#comment-16994717 ]

Tim Allison commented on TIKA-2224:
-----------------------------------

W00t!  Thank you!  I found a couple of handfuls of OneNote files in CommonCrawl.  I'll grab those and share them shortly.

Think we're ready for the merge into master?  If you want to squash all commits to a single commit from you and then open a PR against master, I'm good with that.  I don't need to have my name on some of the intermediate mods I made. :D

> OneNote formats support - Mime Magic and Parser
> -----------------------------------------------
>
>                 Key: TIKA-2224
>                 URL: https://issues.apache.org/jira/browse/TIKA-2224
>             Project: Tika
>          Issue Type: Improvement
>          Components: mime
>    Affects Versions: 1.14
>            Reporter: Nick Burch
>            Priority: Major
>         Attachments: Sample1.json, Sample1.one, note-ssn-test-mmmm.one
>
>
> As raised at http://stackoverflow.com/questions/41272195/onenote-support-for-apache-tika-parsers, we don't have any magic for the OneNote formats. Several years ago we dug out the file format specs (see http://lucene.472066.n3.nabble.com/Tika-OneNote-Support-td4020393.html), but didn't have volunteer energy to implement a parser. However, armed with those specs, we should be able to come up with some mime magic for detection



--
This message was sent by Atlassian Jira
(v8.3.4#803005)