[jira] Commented: (LUCENE-2167) Implement StandardTokenizer with the UAX#29 Standard

classic Classic list List threaded Threaded
1 message Options
Reply | Threaded
Open this post in threaded view
|

[jira] Commented: (LUCENE-2167) Implement StandardTokenizer with the UAX#29 Standard

JIRA jira@apache.org

    [ https://issues.apache.org/jira/browse/LUCENE-2167?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12887741#action_12887741 ]

Steven Rowe commented on LUCENE-2167:
-------------------------------------

I ran it three more times, and it appears that the difference between ClassicTokenizer, UAX29Tokenizer, and the new StandardTokenizer is in the noise:

||Operation||recsPerRun||rec/s||elapsedSec||
|ClassicTokenizer|1262799|665,682.12|1.90|
|ICUTokenizer|1268451|553,666.94|2.29|
|RBBITokenizer|1268451|575,261.25|2.20|
|StandardTokenizer|1268450|658,935.06|1.92|
|UAX29Tokenizer|1268451|642,579.00|1.97|
||Operation||recsPerRun||rec/s||elapsedSec||
|ClassicTokenizer|1262799|668,501.31|1.89|
|ICUTokenizer|1268451|546,275.19|2.32|
|RBBITokenizer|1268451|563,255.31|2.25|
|StandardTokenizer|1268450|651,824.25|1.95|
|UAX29Tokenizer|1268451|664,806.62|1.91|
||Operation||recsPerRun||rec/s||elapsedSec||
|ClassicTokenizer|1262799|674,932.69|1.87|
|ICUTokenizer|1268451|541,841.50|2.34|
|RBBITokenizer|1268451|586,431.38|2.16|
|StandardTokenizer|1268450|635,814.56|2.00|
|UAX29Tokenizer|1268451|650,487.69|1.95|

> Implement StandardTokenizer with the UAX#29 Standard
> ----------------------------------------------------
>
>                 Key: LUCENE-2167
>                 URL: https://issues.apache.org/jira/browse/LUCENE-2167
>             Project: Lucene - Java
>          Issue Type: New Feature
>          Components: contrib/analyzers
>    Affects Versions: 3.1
>            Reporter: Shyamal Prasad
>            Assignee: Robert Muir
>            Priority: Minor
>         Attachments: LUCENE-2167-jflex-tld-macro-gen.patch, LUCENE-2167-jflex-tld-macro-gen.patch, LUCENE-2167-jflex-tld-macro-gen.patch, LUCENE-2167-lucene-buildhelper-maven-plugin.patch, LUCENE-2167.benchmark.patch, LUCENE-2167.benchmark.patch, LUCENE-2167.benchmark.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, LUCENE-2167.patch, standard.zip
>
>   Original Estimate: 0.5h
>  Remaining Estimate: 0.5h
>
> It would be really nice for StandardTokenizer to adhere straight to the standard as much as we can with jflex. Then its name would actually make sense.
> Such a transition would involve renaming the old StandardTokenizer to EuropeanTokenizer, as its javadoc claims:
> bq. This should be a good tokenizer for most European-language documents
> The new StandardTokenizer could then say
> bq. This should be a good tokenizer for most languages.
> All the english/euro-centric stuff like the acronym/company/apostrophe stuff can stay with that EuropeanTokenizer, and it could be used by the european analyzers.

--
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


---------------------------------------------------------------------
To unsubscribe, e-mail: [hidden email]
For additional commands, e-mail: [hidden email]