【问题标题】:Concurrent processing using Stanford CoreNLP (3.5.2)使用 Stanford CoreNLP (3.5.2) 的并发处理
【发布时间】:2015-06-05 21:44:01
【问题描述】:

我在同时注释多个句子时遇到了并发问题。我不清楚是我做错了什么还是 CoreNLP 中存在错误。

我的目标是使用多个并行运行的线程使用管道“tokenize、ssplit、pos、lemma、ner、parse、dcoref”来注释句子。每个线程分配自己的 StanfordCoreNLP 实例,然后将其用于注释。

问题是有时会抛出异常:

java.util.ConcurrentModificationException
	at java.util.ArrayList$Itr.checkForComodification(ArrayList.java:901)
	at java.util.ArrayList$Itr.next(ArrayList.java:851)
	at java.util.Collections$UnmodifiableCollection$1.next(Collections.java:1042)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:463)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.analyzeNode(GrammaticalStructure.java:488)
	at edu.stanford.nlp.trees.GrammaticalStructure.<init>(GrammaticalStructure.java:201)
	at edu.stanford.nlp.trees.EnglishGrammaticalStructure.<init>(EnglishGrammaticalStructure.java:89)
	at edu.stanford.nlp.semgraph.SemanticGraphFactory.makeFromTree(SemanticGraphFactory.java:139)
	at edu.stanford.nlp.pipeline.DeterministicCorefAnnotator.annotate(DeterministicCorefAnnotator.java:89)
	at edu.stanford.nlp.pipeline.AnnotationPipeline.annotate(AnnotationPipeline.java:68)
	at edu.stanford.nlp.pipeline.StanfordCoreNLP.annotate(StanfordCoreNLP.java:412)

我附上了一个应用程序的示例代码,它在我的 Core i3 370M 笔记本电脑(Win 7 64bit,Java 1.8.0.45 64bit)上大约 20 秒内重现了该问题。此应用程序读取识别文本内涵 (RTE) 语料库的 XML 文件,然后使用标准 Java 并发类同时解析所有句子。本地 RTE XML 文件的路径需要作为命令行参数给出。在我的测试中,我在这里使用了公开可用的 XML 文件: http://www.nist.gov/tac/data/RTE/RTE3-DEV-FINAL.tar.gz

package semante.parser.stanford.server;

import java.io.FileInputStream;
import java.io.InputStreamReader;
import java.io.OutputStream;
import java.io.PrintStream;
import java.nio.charset.StandardCharsets;
import java.util.ArrayList;
import java.util.List;
import java.util.Properties;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.Executors;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicInteger;

import javax.xml.bind.JAXBContext;
import javax.xml.bind.Unmarshaller;
import javax.xml.bind.annotation.XmlAccessType;
import javax.xml.bind.annotation.XmlAccessorType;
import javax.xml.bind.annotation.XmlAttribute;
import javax.xml.bind.annotation.XmlElement;
import javax.xml.bind.annotation.XmlRootElement;

import edu.stanford.nlp.pipeline.Annotation;
import edu.stanford.nlp.pipeline.StanfordCoreNLP;

public class StanfordMultiThreadingTest {

	@XmlRootElement(name = "entailment-corpus")
	@XmlAccessorType (XmlAccessType.FIELD)
	public static class Corpus {
		@XmlElement(name = "pair")
		private List<Pair> pairList = new ArrayList<Pair>();

		public void addPair(Pair p) {pairList.add(p);}
		public List<Pair> getPairList() {return pairList;}
	}

	@XmlRootElement(name="pair")
	public static class Pair {

		@XmlAttribute(name = "id")
		String id;

		@XmlAttribute(name = "entailment")
		String entailment;

		@XmlElement(name = "t")
		String t;

		@XmlElement(name = "h")
		String h;

		private Pair() {}

		public Pair(int id, boolean entailment, String t, String h) {
			this();
			this.id = Integer.toString(id);
			this.entailment = entailment ? "YES" : "NO";
			this.t = t;
			this.h = h;
		}

		public String getId() {return id;}
		public String getEntailment() {return entailment;}
		public String getT() {return t;}
		public String getH() {return h;}
	}
	
	class NullStream extends OutputStream {
		@Override 
		public void write(int b) {}
	};

	private Corpus corpus;
	private Unmarshaller unmarshaller;
	private ExecutorService executor;

	public StanfordMultiThreadingTest() throws Exception {
		javax.xml.bind.JAXBContext jaxbCtx = JAXBContext.newInstance(Pair.class,Corpus.class);
		unmarshaller = jaxbCtx.createUnmarshaller();
		executor = Executors.newFixedThreadPool(Runtime.getRuntime().availableProcessors());
	}

	public void readXML(String fileName) throws Exception {
		System.out.println("Reading XML - Started");
		corpus = (Corpus) unmarshaller.unmarshal(new InputStreamReader(new FileInputStream(fileName), StandardCharsets.UTF_8));
		System.out.println("Reading XML - Ended");
	}

	public void parseSentences() throws Exception {
		System.out.println("Parsing - Started");

		// turn pairs into a list of sentences
		List<String> sentences = new ArrayList<String>();
		for (Pair pair : corpus.getPairList()) {
			sentences.add(pair.getT());
			sentences.add(pair.getH());
		}

		// prepare the properties
		final Properties props = new Properties();
		props.put("annotators", "tokenize, ssplit, pos, lemma, ner, parse, dcoref");

		// first run is long since models are loaded
		new StanfordCoreNLP(props);

		// to avoid the CoreNLP initialization prints (e.g. "Adding annotation pos")
		final PrintStream nullPrintStream = new PrintStream(new NullStream());
		PrintStream err = System.err;
		System.setErr(nullPrintStream);

		int totalCount = sentences.size();
		AtomicInteger counter = new AtomicInteger(0);

		// use java concurrency to parallelize the parsing
		for (String sentence : sentences) {
			executor.execute(new Runnable() {
				@Override
				public void run() {
					try {
						StanfordCoreNLP pipeline = new StanfordCoreNLP(props);
						Annotation annotation = new Annotation(sentence);
						pipeline.annotate(annotation);
						if (counter.incrementAndGet() % 20 == 0) {
							System.out.println("Done: " + String.format("%.2f", counter.get()*100/(double)totalCount));
						};
					} catch (Exception e) {
						System.setErr(err);
						e.printStackTrace();
						System.setErr(nullPrintStream);
						executor.shutdownNow();
					}
				}
			});
		}
		executor.shutdown();
		
		System.out.println("Waiting for parsing to end.");		
		executor.awaitTermination(10, TimeUnit.MINUTES);

		System.out.println("Parsing - Ended");
	}

	public static void main(String[] args) throws Exception {
		StanfordMultiThreadingTest smtt = new StanfordMultiThreadingTest();
		smtt.readXML(args[0]);
		smtt.parseSentences();
	}

}

在我尝试查找一些背景信息时,我遇到了来自斯坦福的 Christopher ManningGabor Angeli 给出的答案,这表明斯坦福 CoreNLP 的当代版本应该是线程安全的。然而,最近在 CoreNLP 3.4.1 版上的bug report 描述了一个并发问题。如标题所述,我使用的是 3.5.2 版本。

我不清楚我所面临的问题是由于错误还是由于我使用软件包的方式有问题。如果有更多知识的人可以对此有所了解,我将不胜感激。我希望示例代码对重现问题有用。谢谢!

[1]:

【问题讨论】:

    标签: multithreading concurrency stanford-nlp


    【解决方案1】:

    您是否尝试过使用threads 选项?您可以为单个StanfordCoreNLP 管道指定多个线程,然后它将并行处理句子。

    例如,如果要在 8 个核心上处理句子,请将 threads 选项设置为 8

    Properties props = new Properties();
    props.put("annotators", "tokenize, ssplit, pos, lemma, ner, parse, dcoref");
    props.put("threads", "8")
    StanfordCoreNLP pipeline  = new StanfordCoreNLP(props);
    

    尽管如此,我认为您的解决方案也应该可以工作,我们将检查是否存在一些并发错误,但同时使用此选项可能会解决您的问题。

    【讨论】:

    • 感谢您的建议。我想尝试一下,但我不确定如何使用该界面。假设我设置了'threads'属性,我应该如何并行传递要注释的句子?使用多个使用同一个 StanfordCoreNLP 实例的线程?还是通过与“annotate()”不同的方法一次传递几个句子?谢谢!来电
    • Annotation的构造函数的参数实际上不是一个句子而是整个文档。在sentence 变量中存储几个(甚至所有)句子,并用“\n”分隔它们。还将选项“ssplit.eolonly”设置为“true”,以防止句子拆分器错误地拆分实际句子。解析后,annotation 对象包含一个句子列表,其中每个句子都有解析、pos、lemma 等注释。
    • 谢谢,我试过了。但是,注释由 '\n' 分隔的多个句子的模式存在问题,或者我做错了什么。我能够解析 100 个句子,但不能解析 1000 或 2000 个句子。当输入 1000 或 2000 个句子时,对 annotate() 的调用将无休止地运行。此外,当我用 100 个句子进行测试时,1、2 或 4 个线程(我的硬件有 4 个)之间的性能几乎没有差异。它比使用单线程并一次只用一个句子调用 annotate() 稍慢。
    • 我在这里有一个更新的示例代码:dl.dropboxusercontent.com/u/21642925/… 要运行它,您可以使用 3 个参数: annotationMode - 'together'(一次调用 annotate(),多个句子由 '\ n') 或 'separated' (多次调用 annotate() 每次调用一个句子); coresMode - 单个内核、一半内核数或所有内核; maxSentences - 要解析的最大句子数。如果您能尝试运行此代码并让我知道您是否设法重现这些问题,我将不胜感激。
    • @RajVJain - 什么答案?
    【解决方案2】:

    我遇到了同样的问题,使用最新的 github 修订版(今天)解决了这个问题。所以我认为这是一个CoreNLP问题,从3.5.2开始就解决了。

    另见CoreNLP on Apache Spark

    【讨论】:

    • 感谢您的更新。我会在他们发布新版本时尝试。
    猜你喜欢
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 2013-10-26
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    • 1970-01-01
    相关资源
    最近更新 更多