Which TWO properties in the Read PDF with OCR activity must be correctly configured to ensure the OCR engine processes the document effectively?
Trap 1: Range
The Range property is vital for controlling which pages are parsed. Specifying a range prevents the bot from unnecessarily processing the entire document if only specific pages are needed, reducing execution time and resource consumption. Proper range configuration is essential for handling large PDF files efficiently during automation.
Trap 2: Password
While the Password property exists, it is only required for encrypted documents. It is not a universal requirement for general PDF processing. Including it in a general configuration list is unnecessary, as it is conditional based on the specific security restrictions applied to the target PDF files being processed.
Trap 3: ExtractImages
ExtractImages is not a standard property of the Read PDF with OCR activity. This logic does not exist at the activity level. Attempting to locate this property suggests a misunderstanding of the activity's capability, as image extraction is typically handled by separate, dedicated activities for document parsing.
- A
OCREngine
The OCREngine property is mandatory as it determines the specific engine used for character recognition. Choosing the right engine, like Tesseract, requires proper configuration of the language and whitelist settings to accurately interpret the document text. Without a correctly assigned engine, the activity cannot initiate the recognition process.
- B
Range
Why it fails: The Range property is vital for controlling which pages are parsed. Specifying a range prevents the bot from unnecessarily processing the entire document if only specific pages are needed, reducing execution time and resource consumption. Proper range configuration is essential for handling large PDF files efficiently during automation.
- C
Password
Why it fails: While the Password property exists, it is only required for encrypted documents. It is not a universal requirement for general PDF processing. Including it in a general configuration list is unnecessary, as it is conditional based on the specific security restrictions applied to the target PDF files being processed.
- D
ExtractImages
Why it fails: ExtractImages is not a standard property of the Read PDF with OCR activity. This logic does not exist at the activity level. Attempting to locate this property suggests a misunderstanding of the activity's capability, as image extraction is typically handled by separate, dedicated activities for document parsing.
- E
Timeout
Why it fails: Timeout is used in UI-based activities but is not a native property for the backend processing performed by the Read PDF with OCR activity. This property does not impact the conversion of PDF content to text. Including it as a critical configuration parameter for PDF reading is technically inaccurate.