Java Windows 64位上的Tess4j:多线程上的异常

Java Windows 64位上的Tess4j:多线程上的异常,java,multithreading,tesseract,ghost4j,tess4j,Java,Multithreading,Tesseract,Ghost4j,Tess4j,我在Windows 64位到OCR扫描PDF上使用tesseract 3和Java 8。我遵循并使用了所需DLL的64位版本,并安装了64位Ghostscript 当我使用普通的@test(无参数)运行单元测试时,代码正确运行,因此我想我已经正确安装了所有内容 当我用两个线程并行运行它时(见下文),我得到一个异常 我已经阅读了相关的线程,但是建议使用我正在使用的Tesseract1(我已经尝试了这两种方法) 有什么想法吗 代码如下: // @Test // works @Test(invoca

我在Windows 64位到OCR扫描PDF上使用tesseract 3和Java 8。我遵循并使用了所需DLL的64位版本,并安装了64位Ghostscript

当我使用普通的@test(无参数)运行单元测试时,代码正确运行,因此我想我已经正确安装了所有内容

当我用两个线程并行运行它时(见下文),我得到一个异常

我已经阅读了相关的线程,但是建议使用我正在使用的Tesseract1(我已经尝试了这两种方法)

有什么想法吗

代码如下:

//  @Test // works
@Test(invocationCount = 2, threadPoolSize = 2)
public void testOcr() throws OcrException, TesseractException {
    File scannedPdf = new File(this.getClass().getClassLoader().getResource("scanned.pdf").getFile());
//  Tesseract instance = Tesseract.getInstance();  // JNA Interface Mapping
    Tesseract1 instance = new Tesseract1(); // JNA Direct Mapping
    String str = instance.doOCR(scannedPdf);
    System.out.println("OCR Result: " + str);
}
这是一个例外:

log4j:WARN No appenders could be found for logger (org.ghost4j.Ghostscript).
log4j:WARN Please initialize the log4j system properly.
log4j:WARN See http://logging.apache.org/log4j/1.2/faq.html#noconfig for more info.
Ιουλ 16, 2014 6:22:23 ΜΜ net.sourceforge.vietocr.PdfUtilities convertPdf2Png
SEVERE: Cannot initialize Ghostscript interpreter. Error code is -21
org.ghost4j.GhostscriptException: Cannot initialize Ghostscript interpreter. Error code is -21
    at org.ghost4j.Ghostscript.initialize(Ghostscript.java:365)
    at net.sourceforge.vietocr.PdfUtilities.convertPdf2Png(Unknown Source)
    at net.sourceforge.vietocr.PdfUtilities.convertPdf2Tiff(Unknown Source)
    at net.sourceforge.vietocr.ImageIOHelper.getIIOImageList(Unknown Source)
    at net.sourceforge.tess4j.Tesseract1.doOCR(Unknown Source)
    at net.sourceforge.tess4j.Tesseract1.doOCR(Unknown Source)
    at OcrUtilsTest.testOcr(OcrUtilsTest.java:19)
    at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
    at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)
    at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
    at java.lang.reflect.Method.invoke(Method.java:483)
    at org.testng.internal.MethodInvocationHelper.invokeMethod(MethodInvocationHelper.java:84)
    at org.testng.internal.Invoker.invokeMethod(Invoker.java:714)
    at org.testng.internal.Invoker.invokeTestMethod(Invoker.java:901)
    at org.testng.internal.Invoker.invokeTestMethods(Invoker.java:1231)
    at org.testng.internal.TestMethodWorker.invokeTestMethods(TestMethodWorker.java:127)
    at org.testng.internal.TestMethodWorker.run(TestMethodWorker.java:111)
    at org.testng.internal.thread.ThreadUtil$2.call(ThreadUtil.java:64)
    at java.util.concurrent.FutureTask.run(FutureTask.java:266)
    at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)
    at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)
    at java.lang.Thread.run(Thread.java:745)

java.lang.Error: Invalid memory access
    at com.sun.jna.Native.invokeInt(Native Method)
    at com.sun.jna.Function.invoke(Function.java:383)
    at com.sun.jna.Function.invoke(Function.java:315)
    at com.sun.jna.Library$Handler.invoke(Library.java:212)
    at com.sun.proxy.$Proxy3.gsapi_init_with_args(Unknown Source)
    at org.ghost4j.Ghostscript.initialize(Ghostscript.java:350)
    at net.sourceforge.vietocr.PdfUtilities.convertPdf2Png(Unknown Source)
    at net.sourceforge.vietocr.PdfUtilities.convertPdf2Tiff(Unknown Source)
    at net.sourceforge.vietocr.ImageIOHelper.getIIOImageList(Unknown Source)
    at net.sourceforge.tess4j.Tesseract1.doOCR(Unknown Source)
    at net.sourceforge.tess4j.Tesseract1.doOCR(Unknown Source)
    at OcrUtilsTest.testOcr(OcrUtilsTest.java:19)
    at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
    at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)
    at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
    at java.lang.reflect.Method.invoke(Method.java:483)
    at org.testng.internal.MethodInvocationHelper.invokeMethod(MethodInvocationHelper.java:84)
    at org.testng.internal.Invoker.invokeMethod(Invoker.java:714)
    at org.testng.internal.Invoker.invokeTestMethod(Invoker.java:901)
    at org.testng.internal.Invoker.invokeTestMethods(Invoker.java:1231)
    at org.testng.internal.TestMethodWorker.invokeTestMethods(TestMethodWorker.java:127)
    at org.testng.internal.TestMethodWorker.run(TestMethodWorker.java:111)
    at org.testng.internal.thread.ThreadUtil$2.call(ThreadUtil.java:64)
    at java.util.concurrent.FutureTask.run(FutureTask.java:266)
    at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)
    at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)
    at java.lang.Thread.run(Thread.java:745)
net.sourceforge.tess4j.TesseractException: javax.imageio.IIOException: I/O error reading header!
    at net.sourceforge.tess4j.Tesseract1.doOCR(Unknown Source)
    at net.sourceforge.tess4j.Tesseract1.doOCR(Unknown Source)
    at OcrUtilsTest.testOcr(OcrUtilsTest.java:19)
    at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
    at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:62)
    at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:43)
    at java.lang.reflect.Method.invoke(Method.java:483)
    at org.testng.internal.MethodInvocationHelper.invokeMethod(MethodInvocationHelper.java:84)
    at org.testng.internal.Invoker.invokeMethod(Invoker.java:714)
    at org.testng.internal.Invoker.invokeTestMethod(Invoker.java:901)
    at org.testng.internal.Invoker.invokeTestMethods(Invoker.java:1231)
    at org.testng.internal.TestMethodWorker.invokeTestMethods(TestMethodWorker.java:127)
    at org.testng.internal.TestMethodWorker.run(TestMethodWorker.java:111)
    at org.testng.internal.thread.ThreadUtil$2.call(ThreadUtil.java:64)
    at java.util.concurrent.FutureTask.run(FutureTask.java:266)
    at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1142)
    at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:617)
    at java.lang.Thread.run(Thread.java:745)
Caused by: javax.imageio.IIOException: I/O error reading header!
    at com.sun.media.imageioimpl.plugins.tiff.TIFFImageReader.readHeader(TIFFImageReader.java:224)
    at com.sun.media.imageioimpl.plugins.tiff.TIFFImageReader.locateImage(TIFFImageReader.java:231)
    at com.sun.media.imageioimpl.plugins.tiff.TIFFImageReader.getNumImages(TIFFImageReader.java:279)
    at net.sourceforge.vietocr.ImageIOHelper.getIIOImageList(Unknown Source)
    ... 18 more
Caused by: java.io.EOFException
    at javax.imageio.stream.ImageInputStreamImpl.readShort(ImageInputStreamImpl.java:229)
    at javax.imageio.stream.ImageInputStreamImpl.readUnsignedShort(ImageInputStreamImpl.java:242)
    at com.sun.media.imageioimpl.plugins.tiff.TIFFImageReader.readHeader(TIFFImageReader.java:199)
    ... 21 more

更新:它似乎与相关。

即使扫描PDF,Tesseract本身也只能将图像转换为文本,而不能转换为PDF

在引擎盖下,Tesser4j使用Ghostscript(通过ghost4j)将每个页面转换为单个图像文件,然后将该文件提供给Tesseract进行OCR。它将结果字符串连接成一个字符串,并返回该字符串


出现异常的原因是Tess4j以不支持多线程的方式使用Ghost4j。如前所述,ghost4j确实通过其高级API提供了多线程支持(实际上,它分别运行不同的Ghostscript实例,每个实例都从不同的JVM调用)。然而,Tess4j使用它的低级API,其中可以使用一个Ghostscript实例。

我今天开始遇到这个问题。。。是的。。。这是因为我使用了多个线程。