A birth certificate photographed on a kitchen table
Your agency receives files in every state imaginable. A court bundle as a scanned PDF. A brochure exported as an image-only PDF. A product catalogue in a protected Excel file. A birth certificate photographed at an angle on a kitchen table, with a thumb in the corner. A PowerPoint with text in images.
Before a translator can work in Trados Studio, memoQ or Phrase, someone has to convert each file into something editable, and check that the conversion has not scrambled the text. In many agencies that someone is a project manager who should be doing something else, or a DTP specialist whose time is better spent on layout.
Why file prep is so fiddly
- OCR output breaks lines mid-sentence, which splits segments and ruins matches.
- Tables, columns and headers come out in the wrong order.
- Protected or password files are discovered only when someone tries to open them.
- Text in images is missed entirely, so the word count is wrong.
- Nobody sees the prep work coming, so it is done in a rush on the day of the job.
What it costs
Word counts that are too low, because text in images or scans was not counted, which means underquoted jobs. Translators working on badly segmented files, producing inconsistent output and poorer matches in the translation memory. PM time spent on conversion instead of project management. And deadlines squeezed because preparation took longer than planned.
| File type | Common problem | What the prep queue does |
|---|---|---|
| Scanned PDF | No editable text | OCR with language detection, quality flagged |
| Image-only PDF | Looks editable, is not | Detected on arrival, routed to OCR |
| Text in images | Missed from word count | Detected and listed for a quote adjustment |
| Protected files | Found late | Flagged on arrival so the client can be asked |
| Broken line breaks | Split segments | Joined where sentences run across lines |
How we build file preparation
- Every incoming file is inspected on arrival, whether from the quote form, portal or inbox, and sorted: editable, scanned, image-only, protected, unsupported.
- Scanned and image-only files go through OCR using a service such as Azure AI Document Intelligence or AWS Textract, with the source language set, and a confidence score recorded per page.
- Common clean-up is applied: rejoining broken lines, removing running headers and page numbers from the text flow, and keeping table structure where possible.
- Pages with low OCR confidence, handwriting or stamps are flagged for a person to check, rather than passed silently to a translator.
- Text in images inside editable files is detected and listed, so the PM can decide whether it is in scope and adjust the quote.
- The prepared file is saved as DOCX or another format your CAT tool handles well, with the original kept alongside for reference.
- A prep queue shows every file, its status and anything waiting for a person, so preparation is visible before the job starts.
For certified or legal documents, the translator always works with the original image alongside the prepared text, because stamps, signatures and handwritten notes matter.
What PMs and translators notice
Problem files are found at the quote stage, not on the morning of the deadline. Word counts include what was previously missed. Translators receive cleaner files, with sentences intact, so their CAT tool segments sensibly and matches improve. PMs spend less time converting and more time managing. And DTP specialists are freed for layout work that genuinely needs them.
Do files arrive like this at your agency?
- PMs convert scanned PDFs by hand.
- Word counts have been wrong because of text in images or scans.
- Translators complain about broken segments.
- Protected files are discovered on the day of the job.
- File preparation is invisible until it causes a delay.