33abcde33abcde
First, I suggest you preprocess your image, for example making the dark parts darker, blur it a little. Feel free to experiment until Tesseract stops seeing letters in the filled-in squares.
Second, you have two options:
One, you can enable hOCR output and try to parse the layout of the scanned letters yourself. hOCR is a subset of HTML and it contains coordinates of all recognized words. Try figuring out where the rows and columns are.
Alternatively, try making Tesseract recognise the layout properly, not rotated 90°.
Anyway, this is what I did:
1. I ran the image through ImageMagick:
$ convert CDZjN.png -deskew 40% -contrast-stretch 7%x10% -filter lanczos -resize 250% ooo.png
2. I created a config file t.conf for Tesseract, disabling vertical text detection and English dictionary:
textord_tabfind_vertical_text 0load_system_dawg 0load_freq_dawg 0load_punc_dawg 0load_number_dawg 0load_unambig_dawg 0load_bigram_dawg 0load_fixed_length_dawgs 0
3. I simply ran it:
$ tesseract ooo.png ooo t.conf ; cat ooo.txt Tesseract Open Source OCR Engine v3.02 with Leptonica01ABC-E 26ABCDE02A CDE 27ABCDEo3 BCDE 28ABCDEo4 BCDE 29ABCDEo5 BCDE 30ABCDE06ABCD. 31ABCDE07A-CDE 32ABCDE08ABC.E 33ABCDEo9 BCDE 34ABCDE10A CDE 35ABCDE11ABCD 36ABCDE12ABC E 37ABCDE13ABC E 38ABCDE14ABCD 39ABCDE15 BCDE 40ABCDE1s BCDE 41ABCDE17 BCDE 42ABCDE18ABCD_ 43ABCDE19AB DE 44ABCDE20AB DE 45ABCDE21ABCDE 46ABCDE22ABCDE 47ABCDE23ABCDE 48ABCDE24ABCDE 49ABCDE25ABCDE 50ABCDE
Not perfect, but passable.
本文来自电脑杂谈,转载请注明本文网址:
http://www.pc-fly.com/a/bofangqi/article-43498-1.html
在叙利亚丢了面子
只是在阴沟里不容易被发觉
但是偏偏三哥就是不开窍