【问题标题】:MRUnit does not work with MultipleOutputsMRUnit 不适用于 MultipleOutputs
【发布时间】:2015-04-30 07:12:15
【问题描述】:

当我使用 MultipleOutputs 运行基本 MRUnit 时,出现以下异常:

java.lang.NullPointerException
at org.apache.hadoop.fs.Path.<init>(Path.java:105)
at org.apache.hadoop.fs.Path.<init>(Path.java:94)
at org.apache.hadoop.mapreduce.lib.output.FileOutputFormat.getDefaultWorkFile(FileOutputFormat.java:264)
at org.apache.hadoop.mapreduce.lib.output.TextOutputFormat.getRecordWriter(TextOutputFormat.java:125)
at org.apache.hadoop.mapreduce.lib.output.MultipleOutputs.getRecordWriter(MultipleOutputs.java:405)
at org.apache.hadoop.mapreduce.lib.output.MultipleOutputs.write(MultipleOutputs.java:387)
at com.skobbler.scratch.MOutputReduce.reduce(MOutputReduce.java:45)
at com.skobbler.scratch.MOutputReduce.reduce(MOutputReduce.java:28)
at org.apache.hadoop.mapreduce.Reducer.run(Reducer.java:164)
at org.apache.hadoop.mrunit.mapreduce.ReduceDriver.run(ReduceDriver.java:265)
at org.apache.hadoop.mrunit.mapreduce.ReducePhaseRunner.runReduce(ReducePhaseRunner.java:85)
at org.apache.hadoop.mrunit.mapreduce.MapReduceDriver.run(MapReduceDriver.java:249)
at org.apache.hadoop.mrunit.TestDriver.runTest(TestDriver.java:640)
at org.apache.hadoop.mrunit.TestDriver.runTest(TestDriver.java:627)

发现请求了ma​​pred.output.dir配置,为空。简单输出不会出现此问题。

MRUnit 代码:

    @Test
public void testMultiOutput() throws IOException{
    MapReduceDriver<LongWritable, Text, Text, Text, Text, Text> mapReduceDriver = createMapReduceDrive();
    mapReduceDriver.withInput(new LongWritable(0L), new Text("a,b"));
    mapReduceDriver.withInput(new LongWritable(0L), new Text("a,c"));
    mapReduceDriver.withMultiOutput("foo", new Text("a"), new Text("2"));
    mapReduceDriver.runTest();
}

private MapReduceDriver<LongWritable, Text, Text, Text, Text, Text> createMapReduceDrive() {
    MOutputMap mapper = new MOutputMap();
    MOutputReduce reducer = new MOutputReduce();
    return MapReduceDriver.newMapReduceDriver(mapper, reducer);
}

如何在不指定 hadoop 系统/输出路径的情况下运行测试。

Hadoop 2,MRUnit 1.1.0

【问题讨论】:

  • 我无法重现您的问题。通常我会看到Missing expected outputs for namedOutput ...java.lang.IllegalArgumentException: Named output 'someOutput' not defined 之类的东西,但不会看到NullPointerException。您能否提供更多代码或可以测试的小样本?
  • 我使用了一种解决方法,模拟了 reducer 的发出,以在测试中使用 context.write() 。可能是 windows/hadoop 版本的问题。

标签: hadoop hdfs mrunit


【解决方案1】:

是的,我遇到了这个问题。但是我从它的源代码中找到了解决方案。

TestDriver.java

您可以使用 getConfiguration() 方法获取 JobConfiguration 对象,然后设置 outputdir。

    Configuration conf = mapReduceDriver.getConfiguration();
    conf.set("mapreduce.output.fileoutputformat.outputdir", "aa");

【讨论】:

    【解决方案2】:

    我最近遇到了同样的问题。我之前使用过@RunWith(SpringJUnit4ClassRunner.class) 注释,但根据https://issues.apache.org/jira/browse/MRUNIT-13https://issues.apache.org/jira/browse/MRUNIT-213 的MRUnit JIRA 问题中的cmets,我们需要使用 @RunWith(PowerMockRunner.class) @PrepareForTest(MyMapper.class) 或 @PrepareForTest(MyReducer.class) 运行使用 MultipleOutputs 的测试。

    我希望这可以帮助遇到此问题的其他人。

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2014-08-09
      相关资源
      最近更新 更多