[jira] [Commented] (NUTCH-2720) ROBOTS metatag ignored when capitalized

Previous Topic Next Topic
 
classic Classic list List threaded Threaded
1 message Options
Reply | Threaded
Open this post in threaded view
|

[jira] [Commented] (NUTCH-2720) ROBOTS metatag ignored when capitalized

Hudson (Jira)

    [ https://issues.apache.org/jira/browse/NUTCH-2720?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17129193#comment-17129193 ]

Hudson commented on NUTCH-2720:
-------------------------------

SUCCESS: Integrated in Jenkins build Nutch-trunk #3684 (See [https://builds.apache.org/job/Nutch-trunk/3684/])
NUTCH-2720 ROBOTS metatag ignored when capitalized (snagel: [https://github.com/apache/nutch/commit/f0e1e3d0dc06dadf447d3847f5bc117046e39cd5])
* (edit) src/plugin/parse-tika/src/java/org/apache/nutch/parse/tika/TikaParser.java
NUTCH-2720 ROBOTS metatag ignored when capitalized - move string (snagel: [https://github.com/apache/nutch/commit/aa3a2a617c5c30d40a165025e98c50a0ea839224])
* (edit) src/plugin/parse-html/src/java/org/apache/nutch/parse/html/HTMLMetaProcessor.java
* (edit) src/plugin/parse-tika/src/java/org/apache/nutch/parse/tika/HTMLMetaProcessor.java
* (edit) src/plugin/parse-tika/src/java/org/apache/nutch/parse/tika/TikaParser.java
* (edit) src/java/org/apache/nutch/metadata/Nutch.java
* (edit) src/java/org/apache/nutch/indexer/IndexerMapReduce.java


> ROBOTS metatag ignored when capitalized
> ---------------------------------------
>
>                 Key: NUTCH-2720
>                 URL: https://issues.apache.org/jira/browse/NUTCH-2720
>             Project: Nutch
>          Issue Type: Bug
>          Components: indexer, robots
>    Affects Versions: 1.15
>            Reporter: Felix Zett
>            Assignee: Sebastian Nagel
>            Priority: Minor
>             Fix For: 1.17
>
>         Attachments: noindex.html
>
>
> As discussed [on the mailing list|https://www.mail-archive.com/user@.../msg16516.html], index-metadata fails to ignore a webpage with a capitalized robots metatag such as {{<META NAME="ROBOTS" CONTENT="NOINDEX, FOLLOW">}}. This only applies when parse-tika is used. parse-html will "decapitalize"
> Parsing the attached [^noindex.html] leads to the following results:
> *parse-html:*
> {code:java}
> bin/nutch parsechecker -Dplugin.includes="protocol-httpclient|parse-(html|metatags)|index-metadata" -Dindexer.delete.robots.noindex="true" -Dmetatags.names="robots" -Dindex.parse.md="metatag.robots" http://localhost:8080/noindex.html
> Parse Metadata: [...] metatag.robots=noindex,nofollow robots=noindex,nofollow{code}
> *parse-tika:*
> {code:java}
> bin/nutch parsechecker -Dplugin.includes="protocol-httpclient|parse-(tika|metatags)|index-metadata" -Dindexer.delete.robots.noindex="true" -Dmetatags.names="robots" -Dindex.parse.md="metatag.robots" http://localhost:8080/noindex.html
> Parse Metadata: metatag.robots=NOINDEX,NOFOLLOW  [...] ROBOTS=NOINDEX,NOFOLLOW [...]{code}
>  
> The field being named "ROBOTS" and not "robots" leads to {{parseData.getMeta("robots")}} being {{null}} in [https://github.com/apache/nutch/blob/master/src/java/org/apache/nutch/indexer/IndexerMapReduce.java#L257].



--
This message was sent by Atlassian Jira
(v8.3.4#803005)