Recognize Text in Images with ML Kit on Android

You can use ML Kit to recognize text in images. ML Kit has both a general-purpose API suitable for recognizing text in images, such as the text of a street sign, and an API optimized for recognizing the text of documents. The general-purpose API has both on-device and cloud-based models. Document text recognition is available only as a cloud-based model. See the overview for a comparison of the cloud and on-device models.

Before you begin

If you haven't already, add Firebase to your Android project .

Add the dependencies for the ML Kit Android libraries to your module (app-level) Gradle file (usually app/build.gradle ):

 apply 
  
 plugin 
 : 
  
 ' 
 com 
 . 
 android 
 . 
 application 
 ' 
 apply 
  
 plugin 
 : 
  
 ' 
 com 
 . 
 google 
 . 
 gms 
 . 
 google 
 - 
 services 
 ' 
 dependencies 
  
 { 
  
 // ... 
   
 implementation 
  
 ' 
 com 
 . 
 google 
 . 
 firebase 
 : 
 firebase 
 - 
 ml 
 - 
 vision 
 : 
 24.0.3 
 ' 
 }

Optional but recommended : If you use the on-device API, configure your app to automatically download the ML model to the device after your app is installed from the Play Store.
To do so, add the following declaration to your app's AndroidManifest.xml file:
```
< application 
 ... 
 > 
 ... 
< meta 
 - 
 data 
 android 
 : 
 name 
 = 
 "com.google.firebase.ml.vision.DEPENDENCIES" 
 android 
 : 
 value 
 = 
 "ocr" 
 / 
>
  < ! 
 -- 
 To 
 use 
 multiple 
 models 
 : 
 android 
 : 
 value 
 = 
 "ocr,model2,model3" 
 -- 
>
< / 
 application 
 > 
```
If you do not enable install-time model downloads, the model will be downloaded the first time you run the on-device detector. Requests you make before the download has completed will produce no results.
If you want to use the Cloud-based model, and you have not already enabled the Cloud-based APIs for your project, do so now:
1. Open the ML Kit APIs page of the Firebase console.
2. If you have not already upgraded your project to a Blaze pricing plan, click Upgrade to do so. (You will be prompted to upgrade only if your project isn't on the Blaze plan.)
  
  Only Blaze-level projects can use Cloud-based APIs.
3. If Cloud-based APIs aren't already enabled, click Enable Cloud-based APIs .
Before you deploy to production an app that uses a Cloud API, you should take some additional steps to prevent and mitigate the effect of unauthorized API access .

If you want to use only the on-device model, you can skip this step.

Now you are ready to start recognizing text in images.

Input image guidelines

For ML Kit to accurately recognize text, input images must contain text that is represented by sufficient pixel data. Ideally, for Latin text, each character should be at least 16x16 pixels. For Chinese, Japanese, and Korean text (only supported by the cloud-based APIs), each character should be 24x24 pixels. For all languages, there is generally no accuracy benefit for characters to be larger than 24x24 pixels.

So, for example, a 640x480 image might work well to scan a business card that occupies the full width of the image. To scan a document printed on letter-sized paper, a 720x1280 pixel image might be required.
Poor image focus can hurt text recognition accuracy. If you aren't getting acceptable results, try asking the user to recapture the image.
If you are recognizing text in a real-time application, you might also want to consider the overall dimensions of the input images. Smaller images can be processed faster, so to reduce latency, capture images at lower resolutions (keeping in mind the above accuracy requirements) and ensure that the text occupies as much of the image as possible. Also see Tips to improve real-time performance .

Recognize text in images

To recognize text in an image using either an on-device or cloud-based model, run the text recognizer as described below.

1. Run the text recognizer

To recognize text in an image, create a FirebaseVisionImage object from either a Bitmap , media.Image , ByteBuffer , byte array, or a file on the device. Then, pass the FirebaseVisionImage object to the FirebaseVisionTextRecognizer 's processImage method.

Create a FirebaseVisionImage object from your image.

To create a FirebaseVisionImage object from a media.Image object, such as when capturing an image from a device's camera, pass the media.Image object and the image's rotation to FirebaseVisionImage.fromMediaImage() .

If you use the CameraX library, the OnImageCapturedListener and ImageAnalysis.Analyzer classes calculate the rotation value for you, so you just need to convert the rotation to one of ML Kit's ROTATION_ constants before calling FirebaseVisionImage.fromMediaImage() :

Java

 private 
  
 class 
 YourAnalyzer 
  
 implements 
  
 ImageAnalysis 
 . 
 Analyzer 
  
 { 
  
 private 
  
 int 
  
 degreesToFirebaseRotation 
 ( 
 int 
  
 degrees 
 ) 
  
 { 
  
 switch 
  
 ( 
 degrees 
 ) 
  
 { 
  
 case 
  
 0 
 : 
  
 return 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_0 
 ; 
  
 case 
  
 90 
 : 
  
 return 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_90 
 ; 
  
 case 
  
 180 
 : 
  
 return 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_180 
 ; 
  
 case 
  
 270 
 : 
  
 return 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_270 
 ; 
  
 default 
 : 
  
 throw 
  
 new 
  
 IllegalArgumentException 
 ( 
  
 "Rotation must be 0, 90, 180, or 270." 
 ); 
  
 } 
  
 } 
  
 @Override 
  
 public 
  
 void 
  
 analyze 
 ( 
 ImageProxy 
  
 imageProxy 
 , 
  
 int 
  
 degrees 
 ) 
  
 { 
  
 if 
  
 ( 
 imageProxy 
  
 == 
  
 null 
  
 || 
  
 imageProxy 
 . 
 getImage 
 () 
  
 == 
  
 null 
 ) 
  
 { 
  
 return 
 ; 
  
 } 
  
 Image 
  
 mediaImage 
  
 = 
  
 imageProxy 
 . 
 getImage 
 (); 
  
 int 
  
 rotation 
  
 = 
  
 degreesToFirebaseRotation 
 ( 
 degrees 
 ); 
  
 FirebaseVisionImage 
  
 image 
  
 = 
  
 FirebaseVisionImage 
 . 
 fromMediaImage 
 ( 
 mediaImage 
 , 
  
 rotation 
 ); 
  
 // Pass image to an ML Kit Vision API 
  
 // ... 
  
 } 
 }

Kotlin

 private 
  
 class 
  
 YourImageAnalyzer 
  
 : 
  
 ImageAnalysis 
 . 
 Analyzer 
  
 { 
  
 private 
  
 fun 
  
 degreesToFirebaseRotation 
 ( 
 degrees 
 : 
  
 Int 
 ): 
  
 Int 
  
 = 
  
 when 
 ( 
 degrees 
 ) 
  
 { 
  
 0 
  
 -> 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_0 
  
 90 
  
 -> 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_90 
  
 180 
  
 -> 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_180 
  
 270 
  
 -> 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_270 
  
 else 
  
 -> 
  
 throw 
  
 Exception 
 ( 
 "Rotation must be 0, 90, 180, or 270." 
 ) 
  
 } 
  
 override 
  
 fun 
  
 analyze 
 ( 
 imageProxy 
 : 
  
 ImageProxy?, 
  
 degrees 
 : 
  
 Int 
 ) 
  
 { 
  
 val 
  
 mediaImage 
  
 = 
  
 imageProxy 
 ?. 
 image 
  
 val 
  
 imageRotation 
  
 = 
  
 degreesToFirebaseRotation 
 ( 
 degrees 
 ) 
  
 if 
  
 ( 
 mediaImage 
  
 != 
  
 null 
 ) 
  
 { 
  
 val 
  
 image 
  
 = 
  
 FirebaseVisionImage 
 . 
 fromMediaImage 
 ( 
 mediaImage 
 , 
  
 imageRotation 
 ) 
  
 // Pass image to an ML Kit Vision API 
  
 // ... 
  
 } 
  
 } 
 }

If you don't use a camera library that gives you the image's rotation, you can calculate it from the device's rotation and the orientation of camera sensor in the device:

Java

 private 
  
 static 
  
 final 
  
 SparseIntArray 
  
 ORIENTATIONS 
  
 = 
  
 new 
  
 SparseIntArray 
 (); 
 static 
  
 { 
  
 ORIENTATIONS 
 . 
 append 
 ( 
 Surface 
 . 
 ROTATION_0 
 , 
  
 90 
 ); 
  
 ORIENTATIONS 
 . 
 append 
 ( 
 Surface 
 . 
 ROTATION_90 
 , 
  
 0 
 ); 
  
 ORIENTATIONS 
 . 
 append 
 ( 
 Surface 
 . 
 ROTATION_180 
 , 
  
 270 
 ); 
  
 ORIENTATIONS 
 . 
 append 
 ( 
 Surface 
 . 
 ROTATION_270 
 , 
  
 180 
 ); 
 } 
 /** 
 * Get the angle by which an image must be rotated given the device's current 
 * orientation. 
 */ 
 @RequiresApi 
 ( 
 api 
  
 = 
  
 Build 
 . 
 VERSION_CODES 
 . 
 LOLLIPOP 
 ) 
 private 
  
 int 
  
 getRotationCompensation 
 ( 
 String 
  
 cameraId 
 , 
  
 Activity 
  
 activity 
 , 
  
 Context 
  
 context 
 ) 
  
 throws 
  
 CameraAccessException 
  
 { 
  
 // Get the device's current rotation relative to its "native" orientation. 
  
 // Then, from the ORIENTATIONS table, look up the angle the image must be 
  
 // rotated to compensate for the device's rotation. 
  
 int 
  
 deviceRotation 
  
 = 
  
 activity 
 . 
 getWindowManager 
 (). 
 getDefaultDisplay 
 (). 
 getRotation 
 (); 
  
 int 
  
 rotationCompensation 
  
 = 
  
 ORIENTATIONS 
 . 
 get 
 ( 
 deviceRotation 
 ); 
  
 // On most devices, the sensor orientation is 90 degrees, but for some 
  
 // devices it is 270 degrees. For devices with a sensor orientation of 
  
 // 270, rotate the image an additional 180 ((270 + 270) % 360) degrees. 
  
 CameraManager 
  
 cameraManager 
  
 = 
  
 ( 
 CameraManager 
 ) 
  
 context 
 . 
 getSystemService 
 ( 
 CAMERA_SERVICE 
 ); 
  
 int 
  
 sensorOrientation 
  
 = 
  
 cameraManager 
  
 . 
 getCameraCharacteristics 
 ( 
 cameraId 
 ) 
  
 . 
 get 
 ( 
 CameraCharacteristics 
 . 
 SENSOR_ORIENTATION 
 ); 
  
 rotationCompensation 
  
 = 
  
 ( 
 rotationCompensation 
  
 + 
  
 sensorOrientation 
  
 + 
  
 270 
 ) 
  
 % 
  
 360 
 ; 
  
 // Return the corresponding FirebaseVisionImageMetadata rotation value. 
  
 int 
  
 result 
 ; 
  
 switch 
  
 ( 
 rotationCompensation 
 ) 
  
 { 
  
 case 
  
 0 
 : 
  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_0 
 ; 
  
 break 
 ; 
  
 case 
  
 90 
 : 
  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_90 
 ; 
  
 break 
 ; 
  
 case 
  
 180 
 : 
  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_180 
 ; 
  
 break 
 ; 
  
 case 
  
 270 
 : 
  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_270 
 ; 
  
 break 
 ; 
  
 default 
 : 
  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_0 
 ; 
  
 Log 
 . 
 e 
 ( 
 TAG 
 , 
  
 "Bad rotation value: " 
  
 + 
  
 rotationCompensation 
 ); 
  
 } 
  
 return 
  
 result 
 ; 
 } 
   VisionImage 
 . 
 java

Kotlin

 private 
  
 val 
  
 ORIENTATIONS 
  
 = 
  
 SparseIntArray 
 () 
 init 
  
 { 
  
 ORIENTATIONS 
 . 
 append 
 ( 
 Surface 
 . 
 ROTATION_0 
 , 
  
 90 
 ) 
  
 ORIENTATIONS 
 . 
 append 
 ( 
 Surface 
 . 
 ROTATION_90 
 , 
  
 0 
 ) 
  
 ORIENTATIONS 
 . 
 append 
 ( 
 Surface 
 . 
 ROTATION_180 
 , 
  
 270 
 ) 
  
 ORIENTATIONS 
 . 
 append 
 ( 
 Surface 
 . 
 ROTATION_270 
 , 
  
 180 
 ) 
 } 
 /** 
 * Get the angle by which an image must be rotated given the device's current 
 * orientation. 
 */ 
 @RequiresApi 
 ( 
 api 
  
 = 
  
 Build 
 . 
 VERSION_CODES 
 . 
 LOLLIPOP 
 ) 
 @Throws 
 ( 
 CameraAccessException 
 :: 
 class 
 ) 
 private 
  
 fun 
  
 getRotationCompensation 
 ( 
 cameraId 
 : 
  
 String 
 , 
  
 activity 
 : 
  
 Activity 
 , 
  
 context 
 : 
  
 Context 
 ): 
  
 Int 
  
 { 
  
 // Get the device's current rotation relative to its "native" orientation. 
  
 // Then, from the ORIENTATIONS table, look up the angle the image must be 
  
 // rotated to compensate for the device's rotation. 
  
 val 
  
 deviceRotation 
  
 = 
  
 activity 
 . 
 windowManager 
 . 
 defaultDisplay 
 . 
 rotation 
  
 var 
  
 rotationCompensation 
  
 = 
  
 ORIENTATIONS 
 . 
 get 
 ( 
 deviceRotation 
 ) 
  
 // On most devices, the sensor orientation is 90 degrees, but for some 
  
 // devices it is 270 degrees. For devices with a sensor orientation of 
  
 // 270, rotate the image an additional 180 ((270 + 270) % 360) degrees. 
  
 val 
  
 cameraManager 
  
 = 
  
 context 
 . 
 getSystemService 
 ( 
 CAMERA_SERVICE 
 ) 
  
 as 
  
 CameraManager 
  
 val 
  
 sensorOrientation 
  
 = 
  
 cameraManager 
  
 . 
 getCameraCharacteristics 
 ( 
 cameraId 
 ) 
  
 . 
 get 
 ( 
 CameraCharacteristics 
 . 
 SENSOR_ORIENTATION 
 ) 
 !! 
  
 rotationCompensation 
  
 = 
  
 ( 
 rotationCompensation 
  
 + 
  
 sensorOrientation 
  
 + 
  
 270 
 ) 
  
 % 
  
 360 
  
 // Return the corresponding FirebaseVisionImageMetadata rotation value. 
  
 val 
  
 result 
 : 
  
 Int 
  
 when 
  
 ( 
 rotationCompensation 
 ) 
  
 { 
  
 0 
  
 - 
>  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_0 
  
 90 
  
 - 
>  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_90 
  
 180 
  
 - 
>  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_180 
  
 270 
  
 - 
>  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_270 
  
 else 
  
 - 
>  
 { 
  
 result 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 ROTATION_0 
  
 Log 
 . 
 e 
 ( 
 TAG 
 , 
  
 "Bad rotation value: 
 $ 
 rotationCompensation 
 " 
 ) 
  
 } 
  
 } 
  
 return 
  
 result 
 } 
   VisionImage 
 . 
 kt

Then, pass the media.Image object and the rotation value to FirebaseVisionImage.fromMediaImage() :

Java

 FirebaseVisionImage 
  
 image 
  
 = 
  
 FirebaseVisionImage 
 . 
 fromMediaImage 
 ( 
 mediaImage 
 , 
  
 rotation 
 ); 
   VisionImage 
 . 
 java

Kotlin

 val 
  
 image 
  
 = 
  
 FirebaseVisionImage 
 . 
 fromMediaImage 
 ( 
 mediaImage 
 , 
  
 rotation 
 ) 
   VisionImage 
 . 
 kt

To create a FirebaseVisionImage object from a file URI, pass the app context and file URI to FirebaseVisionImage.fromFilePath() . This is useful when you use an ACTION_GET_CONTENT intent to prompt the user to select an image from their gallery app.

Java

 FirebaseVisionImage 
  
 image 
 ; 
 try 
  
 { 
  
 image 
  
 = 
  
 FirebaseVisionImage 
 . 
 fromFilePath 
 ( 
 context 
 , 
  
 uri 
 ); 
 } 
  
 catch 
  
 ( 
 IOException 
  
 e 
 ) 
  
 { 
  
 e 
 . 
 printStackTrace 
 (); 
 } 
   VisionImage 
 . 
 java

Kotlin

 val 
  
 image 
 : 
  
 FirebaseVisionImage 
 try 
  
 { 
  
 image 
  
 = 
  
 FirebaseVisionImage 
 . 
 fromFilePath 
 ( 
 context 
 , 
  
 uri 
 ) 
 } 
  
 catch 
  
 ( 
 e 
 : 
  
 IOException 
 ) 
  
 { 
  
 e 
 . 
 printStackTrace 
 () 
 } 
   VisionImage 
 . 
 kt

To create a FirebaseVisionImage object from a ByteBuffer or a byte array, first calculate the image rotation as described above for media.Image input.

Then, create a FirebaseVisionImageMetadata object that contains the image's height, width, color encoding format, and rotation:

Java

 FirebaseVisionImageMetadata 
  
 metadata 
  
 = 
  
 new 
  
 FirebaseVisionImageMetadata 
 . 
 Builder 
 () 
  
 . 
 setWidth 
 ( 
 480 
 ) 
  
 // 480x360 is typically sufficient for 
  
 . 
 setHeight 
 ( 
 360 
 ) 
  
 // image recognition 
  
 . 
 setFormat 
 ( 
 FirebaseVisionImageMetadata 
 . 
 IMAGE_FORMAT_NV21 
 ) 
  
 . 
 setRotation 
 ( 
 rotation 
 ) 
  
 . 
 build 
 (); 
   VisionImage 
 . 
 java

Kotlin

 val 
  
 metadata 
  
 = 
  
 FirebaseVisionImageMetadata 
 . 
 Builder 
 () 
  
 . 
 setWidth 
 ( 
 480 
 ) 
  
 // 480x360 is typically sufficient for 
  
 . 
 setHeight 
 ( 
 360 
 ) 
  
 // image recognition 
  
 . 
 setFormat 
 ( 
 FirebaseVisionImageMetadata 
 . 
 IMAGE_FORMAT_NV21 
 ) 
  
 . 
 setRotation 
 ( 
 rotation 
 ) 
  
 . 
 build 
 () 
   VisionImage 
 . 
 kt

Use the buffer or array, and the metadata object, to create a FirebaseVisionImage object:

Java

 FirebaseVisionImage 
  
 image 
  
 = 
  
 FirebaseVisionImage 
 . 
 fromByteBuffer 
 ( 
 buffer 
 , 
  
 metadata 
 ); 
 // Or: FirebaseVisionImage image = FirebaseVisionImage.fromByteArray(byteArray, metadata);  VisionImage.java

Recognize Text in Images with ML Kit on Android Stay organized with collections Save and categorize content based on your preferences.

Before you begin

Input image guidelines

Recognize text in images

1. Run the text recognizer

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

2. Extract text from blocks of recognized text

Java

Kotlin

Tips to improve real-time performance

Next steps

Recognize text in images of documents

1. Run the text recognizer

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

Java

Kotlin

2. Extract text from blocks of recognized text

Java

Kotlin

Next steps

Recognize Text in Images with ML Kit on Android