Apache pig 迭代连接集后出现清管器错误1066。

Apache pig 迭代连接集后出现清管器错误1066。,apache-pig,Apache Pig,正在尝试将一个月中有天数的集合与年-月键上的数据集合并。在我加入并尝试在集合上执行FOREACH之后,我得到一个错误:1066。。。后端错误:标量在输出中有多行 以下是一个有相同问题的缩写集: $ hadoop fs -cat DIM/\* 2011,01,31 2011,02,28 2011,03,31 2011,04,30 2011,05,31 2011,06,30 2011,07,31 2011,08,31 2011,09,30 2011,10,31 2011,11,30 2011,12,

正在尝试将一个月中有天数的集合与年-月键上的数据集合并。在我加入并尝试在集合上执行FOREACH之后,我得到一个错误:1066。。。后端错误:标量在输出中有多行

以下是一个有相同问题的缩写集:

$ hadoop fs -cat DIM/\*
2011,01,31
2011,02,28
2011,03,31
2011,04,30
2011,05,31
2011,06,30
2011,07,31
2011,08,31
2011,09,30
2011,10,31
2011,11,30
2011,12,31

$ hadoop fs -cat ACCT/\*
2011,7,26,key1,23.25,2470.0
2011,7,26,key2,10.416666666666668,232274.08333333334
2011,7,26,key3,82.83333333333333,541377.25
2011,7,26,key4,78.5,492823.33333333326
2011,7,26,key5,110.83333333333334,729811.9166666667
2011,7,26,key6,102.16666666666666,675941.25
2011,7,26,key7,118.91666666666666,770896.75
然后咕哝着说:

grunt> DIM = LOAD 'DIM' USING PigStorage(',') AS (year:int, month:int, days:int);
grunt> ACCT = LOAD 'ACCT' USING PigStorage(',') AS (year:int, month:int, day: int, account:chararray, metric1:double, metric2:double);
grunt> AjD = JOIN ACCT BY (year,month), DIM  BY (year,month) USING 'replicated';
grunt> dump AjD;
...
(2011,7,26,key1,23.25,2470.0,2011,7,31)
(2011,7,26,key2,10.416666666666668,232274.08333333334,2011,7,31)
(2011,7,26,key3,82.83333333333333,541377.25,2011,7,31)
(2011,7,26,key4,78.5,492823.33333333326,2011,7,31)
(2011,7,26,key5,110.83333333333334,729811.9166666667,2011,7,31)
(2011,7,26,key6,102.16666666666666,675941.25,2011,7,31)
(2011,7,26,key7,118.91666666666666,770896.75,2011,7,31)
grunt> describe AjD;
AjD: {ACCT::year: int,ACCT::month: int,ACCT::day: int,ACCT::account: chararray,ACCT::metric1: double,ACCT::metric2: double,DIM::year: int,DIM::month: int,DIM::days: int}

grunt> FINAL = FOREACH AjD
>> GENERATE ACCT.year, ACCT.month, ACCT.account, (ACCT.metric2 / DIM.days);
grunt> dump FINAL;
...
ERROR org.apache.pig.tools.grunt.Grunt - ERROR 1066: Unable to open iterator for alias FINAL. Backend error : Scalar has more than one row in the output. 1st : (2011,7,26,key1,23.25,2470.0), 2nd :(2011,7,26,key2,10.416666666666668,232274.08333333334)
但是,如果我存储它并重新加载它以摆脱“join”模式,它就会工作:

grunt> STORE AjD INTO 'AjD' using PigStorage(',');
grunt> AjD2 = LOAD 'AjD' USING PigStorage(',') AS (year:int, month:int, day:int, account:chararray, metric1:double, metric2:double, year2:int, month2:int, days:int);

grunt> FINAL = FOREACH AjD2                                                                   
>> GENERATE year, month, account, (metric2 /days);         

grunt> dump FINAL;
...
(2011,7,key1,79.6774193548387)
(2011,7,key2,7492.712365591398)
(2011,7,key3,17463.782258064515)
(2011,7,key4,15897.526881720427)
(2011,7,key5,23542.319892473122)
(2011,7,key6,21804.5564516129)
(2011,7,key7,24867.637096774193)
有没有一种方法可以在不存储和重新加载的情况下对联接集进行迭代(FOREACH)?

您是否尝试过使用指定获取哪一列的方法

(ACCT.metric2/DIM.days)
替换为
(ACCT::metric2/DIM::days)

e、 g

您是否尝试过使用指定要获取的列的名称

(ACCT.metric2/DIM.days)
替换为
(ACCT::metric2/DIM::days)

e、 g


所有列限定符都必须是“:”,谢谢您的回答。添加一个链接。@shoover通过链接和反向链接这两个问题,创建了一个无限循环。:)所有列限定符都必须是“:”,谢谢您的回答。添加一个链接。@shoover通过链接和反向链接这两个问题,创建了一个无限循环。:)对于在这里查找时找到此帖子的人是。对于在这里查找时找到此帖子的人是。
...
FINAL = FOREACH AjD
        GENERATE
             ACCT.year, ACCT.month, ACCT.account,(ACCT::metric2 / DIM::days);