【发布时间】:2014-12-02 05:05:11
【问题描述】:
我正在尝试使用 listOfWords 文件来仅计算任何输入文件中的那些单词。即使我已经验证该文件在 HDFS 中的正确位置,也会出现 FileNotFound 错误。
内部驱动:
Configuration conf = new Configuration();
DistributedCache.addCacheFile(new URI("/user/training/listOfWords"), conf);
Job job = new Job(conf,"CountEachWord Job");
内部映射器:
private Path[] ref_file;
ArrayList<String> globalList = new ArrayList<String>();
public void setup(Context context) throws IOException{
this.ref_file = DistributedCache.getLocalCacheFiles(context.getConfiguration());
FileSystem fs = FileSystem.get(context.getConfiguration());
FSDataInputStream in_file = fs.open(ref_file[0]);
System.out.println("File opened");
BufferedReader br = new BufferedReader(new InputStreamReader(in_file));//each line of reference file
System.out.println("BufferReader invoked");
String eachLine = null;
while((eachLine = br.readLine()) != null)
{
System.out.println("eachLine is: "+ eachLine);
globalList.add(eachLine);
}
}
错误信息:
hadoop jar CountOnlyMatchWords.jar CountEachWordDriver Rhymes CountMatchWordsOut1
Warning: $HADOOP_HOME is deprecated.
14/10/07 22:28:59 WARN mapred.JobClient: Use GenericOptionsParser for parsing the arguments. Applications should implement Tool for the same.
14/10/07 22:28:59 INFO input.FileInputFormat: Total input paths to process : 1
14/10/07 22:28:59 INFO util.NativeCodeLoader: Loaded the native-hadoop library
14/10/07 22:28:59 WARN snappy.LoadSnappy: Snappy native library not loaded
14/10/07 22:29:00 INFO mapred.JobClient: Running job: job_201409300531_0041
14/10/07 22:29:01 INFO mapred.JobClient: map 0% reduce 0%
14/10/07 22:29:14 INFO mapred.JobClient: Task Id : attempt_201409300531_0041_m_000000_0, Status : FAILED
java.io.FileNotFoundException: File does not exist: /home/training/hadoop-temp/mapred/local /taskTracker/distcache/5910352135771601888_2043607380_1633197895/localhost/user/training/listOfWords
我已验证上述文件存在于 HDFS 中。我也尝试使用 localRunner。仍然没有工作。
【问题讨论】:
-
代替 DistributedCache.addCacheFile(new URI("/user/training/listOfWords"), conf);试试这个 DistributedCache.addCacheFile(new URI("/user/training/listOfWords"), job.getConfiguration());
标签: java hadoop mapreduce distributed-caching