Tesseract與ios的集成

最近接觸了一個關於圖文識別的項目,項目組決定使用Tesseract,大概查了一下可是對於ios的支持貌似不是很好,如下是爲Tesseract包裝過的oc版本。php

 

原文連接:https://github.com/ldiqual/tesseract-ioshtml

 

Tesseract for iOS

tesseract-ios is not actively maintained anymore. I encourage you to use gali8's Tesseract-OCR-iOS instead.

About

Tesseract-ios is an Objective-C wrapper for Tesseract OCR.ios

This project couldn't exist without the Ângelo Suzuki's blog post. A lot of code came from his article.c++

Requirements

  • iOS SDK 6.0, iOS 5.0+ (there is no support for armv6)
  • Tesseract and Leptonica libraries from the tesseract-ios-lib repo.

Installation

  • Add tesseract-ios as a group, and tessdata by reference to your project:

  • Go to your project settings, and ensure that C++ Standard Library => libstdc++:

Usage

Here is the default workflow to extract text from an image:git

  • Instantiate Tesseract with data path and language
  • Set variables (character set, …)
  • Set the image to analyze
  • Start recognition
  • Get recognized text
  • Clear

Code Sample

#import "Tesseract.h"

Tesseract* tesseract = [[Tesseract alloc] initWithDataPath:@"tessdata" language:@"eng"];
[tesseract setVariableValue:@"0123456789" forKey:@"tessedit_char_whitelist"];
[tesseract setImage:[UIImage imageNamed:@"image_sample.jpg"]];
[tesseract recognize];

NSLog(@"%@", [tesseract recognizedText]);
[tesseract clear];

Method reference

-initWithDataPath:language:

- (id)initWithDataPath:(NSString *)dataPath language:(NSString *)languagegithub

Initialize a new Tesseract instance.web

  • dataPath: a relative path from the application bundle to the .traineddata files. You can find these files from the tesseract downloads section.
  • language: language used for recognition. Ex: eng. Tesseract will search for a eng.traineddatafile in the dataPath directory.

Returns nil if instanciation failed.app

-setVariableValue:forKey:

- (void)setVariableValue:(NSString *)value forKey:(NSString *)keyide

Set Tesseract variable key to value. See http://www.sk-spell.sk.cx/tesseract-ocr-en-variables for a complete (but not up-to-date) list.wordpress

For instance, use tessedit_char_whitelist to restrict characters to a specific set.

-setImage:

- (void)setImage:(UIImage *)image

Set the image to recognize.

-setLanguage:

- (BOOL)setLanguage:(NSString *)language

Override the language defined with -initWithDataPath:language:.

-recognize

- (BOOL)recognize

Start text recognition. You might want to launch this process in background with NSObject's -performSelectorInBackground:withObject:.

-recognizedText

- (NSString *)recognizedText

Get the text extracted from the image.

-clear

- (void) clear

Clears Tesseract object after text has been recognized from image. Preventing memory leaks.

 

 

備忘:看過一篇博文提到:爲了提升效率須要對圖片進行預處理(二值化、灰度、傾斜校訂和圖片切割),傾斜校訂和圖片切割能夠用openCV的庫處理

博文連接:http://www.cocoachina.com/bbs/read.php?tid=123463 (該博文有demo這裏就很少鏈了)

 

比較有用的連接:

相關文章
相關標籤/搜索