Warning: file_get_contents(/data/phpspider/zhask/data//catemap/2/python/289.json): failed to open stream: No such file or directory in /data/phpspider/zhask/libs/function.php on line 167

Warning: Invalid argument supplied for foreach() in /data/phpspider/zhask/libs/tag.function.php on line 1116

Notice: Undefined index: in /data/phpspider/zhask/libs/function.php on line 180

Warning: array_chunk() expects parameter 1 to be array, null given in /data/phpspider/zhask/libs/function.php on line 181

Warning: file_get_contents(/data/phpspider/zhask/data//catemap/1/angularjs/23.json): failed to open stream: No such file or directory in /data/phpspider/zhask/libs/function.php on line 167

Warning: Invalid argument supplied for foreach() in /data/phpspider/zhask/libs/tag.function.php on line 1116

Notice: Undefined index: in /data/phpspider/zhask/libs/function.php on line 180

Warning: array_chunk() expects parameter 1 to be array, null given in /data/phpspider/zhask/libs/function.php on line 181
Python NLTK Brill标记器拆分单词_Python_Nltk_Pos Tagger - Fatal编程技术网

Python NLTK Brill标记器拆分单词

Python NLTK Brill标记器拆分单词,python,nltk,pos-tagger,Python,Nltk,Pos Tagger,我正在使用python版本3.4.1和NLTK版本3,我正在尝试使用它们的Brill标记器 以下是brill标记器的培训代码: import nltk from nltk.tag.brill import * import nltk.tag.brill_trainer as bt from nltk.corpus import brown Template._cleartemplates() templates = fntbl37() tagged_sentences = brown.tagg

我正在使用python版本3.4.1和NLTK版本3,我正在尝试使用它们的Brill标记器

以下是brill标记器的培训代码:

import nltk
from nltk.tag.brill import *
import nltk.tag.brill_trainer as bt
from nltk.corpus import brown

Template._cleartemplates()
templates = fntbl37()
tagged_sentences = brown.tagged_sents(categories = 'news')
tagged_sentences = tagged_sentences[:]
tagger = nltk.tag.BigramTagger(tagged_sentences)
tagger = bt.BrillTaggerTrainer(tagger, templates, trace=3)
tagger = tagger.train(tagged_sentences, max_rules=250)
print(tagger.evaluate(brown.tagged_sents(categories='fiction')[:]))
print(tagger.tag("Hi I am Harry Potter."))
但是,最后一个命令的输出为:

[('H', 'NN'), ('i', 'NN'), (' ', 'NN'), ('I', 'NN'), (' ', 'NN'), ('a', 'AT'), ('m', 'NN'), (' ', 'NN'), ('H', 'NN'), ('a', 'AT'), ('r', 'NN'), ('r', 'NN'), ('y', 'NN'), (' ', 'NN'), ('P', 'NN'), ('o', 'NN'), ('t', 'NN'), ('t', 'NN'), ('e', 'NN'), ('r', 'NN'), ('.', '.')]
如何阻止它将单词拆分为字母并标记字母而不是单词?

标记
Tag()
函数需要一个标记列表作为输入。 因为您给它一个字符串作为输入,所以这个字符串被解释为一个列表。 将字符串转换为列表将提供一个字符列表:

>>> list("abc")
['a', 'b', 'c']
您只需在标记之前将字符串转换为标记列表。例如,使用或仅通过在空白处拆分:

>>> import nltk
>>> nltk.word_tokenize("Hi I am Harry Potter.")
['Hi', 'I', 'am', 'Harry', 'Potter', '.']
>>> "Hi I am Harry Potter.".split(' ')
['Hi', 'I', 'am', 'Harry', 'Potter.']
在标记中添加标记会产生以下结果:

print(tagger.tag(nltk.word_tokenize("Hi I am Harry Potter.")))
[('Hi', 'NN'), ('I', 'PPSS'), ('am', 'VB'), ('Harry', 'NN'), ('Potter', 'NN'), ('.', '.')]