2

I'm using Tesseract but I don't know whether it neglects any nontext area and targets text only. Do I have to remove any nontext area as a preprocessing step for better output?

rmtheis
  • 6,552
  • 11
  • 58
  • 75
chostDevil
  • 1,001
  • 5
  • 16
  • 24

1 Answers1

2

Tesseract has a pretty good algorithm to detect text, but it will eventually give false-positive matches.

Ideally, you would pre-process the image before submitting it to tesseract. Some time ago I engaged in a similar task, so I suggest you take a look at the following material:

Community
  • 1
  • 1
karlphillip
  • 89,883
  • 35
  • 240
  • 408