【问题标题】:How to set output writer in MapReduce如何在 MapReduce 中设置输出编写器
【发布时间】:2011-09-09 15:09:52
【问题描述】:

我正在尝试 (http://code.google.com/p/appengine-mapreduce/) 的 mapreduce 框架,并稍微修改了演示应用程序(使用 mapreduce.input_readers.DatastoreInputReader 而不是 mapreduce.input_readers.BlobstoreZipInputReader)。

我设置了 2 个管道类:

class IndexPipeline(base_handler.PipelineBase):
def run(self):
    output = yield mapreduce_pipeline.MapreducePipeline(
        "index",
        "main.index_map", #added higher up in code
        "main.index_reduce", #added higher up in code
        "mapreduce.input_readers.DatastoreInputReader",
        mapper_params={
            "entity_kind": "model.SearchRecords",
        },
        shards=16)
    yield StoreOutput("Index", output)

class StoreOutput(base_handler.PipelineBase):
    def run(self, mr_type, encoded_key):
        logging.info("output is %s %s" % (mr_type, str(encoded_key)))
        if encoded_key:
            key = db.Key(encoded=encoded_key)
            m = db.get(key)

            yield op.db.Put(m)

然后运行它:

pipeline = IndexPipeline()
pipeline.start()

但我不断收到此错误:

Handler yielded two: ['a'] , but no output writer is set.

我试图在source 中的某个位置设置输出编写器,但没有成功。我发现的唯一一件事是应该在某处设置output_writer_class

有人知道怎么设置吗?

附带说明,StoreOutput 中的 encoded_key 参数似乎始终为 None。

【问题讨论】:

    标签: google-app-engine mapreduce


    【解决方案1】:

    输出 writer 必须定义为 mapreduce_pipeline.MapreducePipeline 的参数(参见 docstring):

    class MapreducePipeline(base_handler.PipelineBase):
      """Pipeline to execute MapReduce jobs.
    
      Args:
        job_name: job name as string.
        mapper_spec: specification of mapper to use.
        reducer_spec: specification of reducer to use.
        input_reader_spec: specification of input reader to read data from.
        output_writer_spec: specification of output writer to save reduce output to.**
        mapper_params: parameters to use for mapper phase.
        reducer_params: parameters to use for reduce phase.
        shards: number of shards to use as int.
        combiner_spec: Optional. Specification of a combine function. If not
          supplied, no combine step will take place. The combine function takes a
          key, list of values and list of previously combined results. It yields
          combined values that might be processed by another combiner call, but will
          eventually end up in reducer. The combiner output key is assumed to be the
          same as the input key.
    
      Returns:
        filenames from output writer.
      """
    

    【讨论】:

      猜你喜欢
      • 1970-01-01
      • 1970-01-01
      • 1970-01-01
      • 2012-12-02
      • 2016-03-04
      • 2018-05-21
      • 1970-01-01
      • 2017-05-20
      相关资源
      最近更新 更多