PDF transcription: getting the text out of a PDF
"Transcribing a PDF" covers two very different jobs. Sometimes the document already contains its text and you simply cannot see it — then you are copying. Sometimes the PDF is a scan, a picture of a page, and the characters have to be recognised first — then you need OCR. Nearly every web service offering either one sends your file to a server, which is why this site will not offer one; it also has no text extractor and no OCR engine, so there is no button here that does it. Here is what actually works on your own machine.
One word of warning before the rest: in Japanese 「文字起こし」 usually means transcribing a recording. This site has no audio or video tools of any kind, so if you want a meeting or a video turned into text, none of the below is for you — the guide is about a PDF you are already holding.
First: does this PDF contain text, or a picture of text?
Open it in your browser (drag the file onto the window) and try to select a line of words with the mouse.
- The text highlights and can be copied — the characters are in the file. You only need a viewer, not OCR.
- Nothing highlights and copying gives nothing — the page is an image, most likely a scan or a photo of a document. Nothing can be copied out, because there are no characters to copy.
A few things can make real text look absent: text drawn as curves or outlines, a scanned page that was then compressed flat, or a file that only opens in a specific reader. If the file opens as a blank page, it is also worth checking the viewer's page count — a one-page document that renders as empty is usually a scan seen from the wrong side, not a broken file.
If the text is there: copy it out
One or two pages, nothing to install.
- Open the PDF in Chrome, Edge, Safari or Firefox.
- Select everything with
Ctrl+A(Cmd+Aon macOS), or drag across the part you need. - Copy with
Ctrl+Cand paste into a text editor, Word or Pages. Pasting as plain text drops the formatting and leaves plain lines of text; columns come out left to right, which is not the reading order for vertical Japanese.
A whole document: pdftotext from Poppler. It is free, it runs on your own machine, and the document is never uploaded:
pdftotext -layout report.pdf report.txt— writes the text to a file.-layoutkeeps the original column positions, which is what you want for tables and forms.pdftotext -f 3 -l 6 report.pdf part.txt— only pages 3 to 6.-rawinstead of-layoutgives reading order without trying to keep the spacing.
One thing worth knowing: a PDF stores text as codes for glyphs, not as readable characters, and the tool needs the font's character map to turn them back into words. Files produced by scanners, by some older printers and by a few office suites are missing that map, and then the extracted text comes out empty or as nonsense. That is a property of the file, not a bug in the tool.
If it is a scan: OCR is the missing step
OCR (optical character recognition) reads the picture of a word and decides which character it is. It is the only way to get words out of a scan, and it is a separate program rather than something a browser does for free.
- Tesseract (free, open source) reads one image at a time:
tesseract page.png out -l jpnwritesout.txt. Japanese needs the language data pack, which is what-l jpnselects. - OCRmyPDF (free, needs Tesseract and Ghostscript) does the whole document and puts the recognised text back on top of the page images as an invisible layer:
ocrmypdf --language jpn scan.pdf searchable.pdf. The result looks identical, but its text can now be selected and searched, which is what most people actually wanted.
Always read the result. OCR is guessing from shapes, so it is a draft, never a transcript to file without checking. The guesses that go wrong most often:
- Vertical Japanese text — the hardest case, and it is the normal case for a Japanese document.
- Furigana (ruby), which is small and gets read as separate characters or dropped.
- Similar-looking kanji and kana, and digits that turn into letters or the other way round.
- Skewed, low-contrast or small print, and anything crossing a fold or a shadow in a photographed page.
- Rules, borders and table lines, which often turn into stray characters.
Uploading is the part to think about
It is worth being clear about where the work happens. Browser OCR, phone camera text recognition and the web services that offer "PDF to text" send the document or the page image to a server to be processed. Local tools — your own viewer, pdftotext, Tesseract, OCRmyPDF — keep the file on your machine, which matters when the document is a contract, a medical record, a payslip or anything with somebody else's personal details in it.
Getting a scan ready so OCR reads it well
This is where the tools on this site do help, because they all work on your own machine:
- Read the pages at 300 DPI. OCR reads pictures, and small characters are the main reason recognition fails. This site cannot render PDF pages itself, so PDF to JPG explains the free local renderers (Ghostscript, Poppler, ImageMagick) that produce them.
- Strip the camera data. Remove EXIF & GPS removes location and camera details from scan images before they go anywhere that keeps them.
- Make the text big enough. Resize to exact pixels scales a page image up when a scan was made at a low resolution.
- Change the format if a tool insists. Convert image format turns a page into PNG when a scanner's output needs changing.
- Trim the file first. Extract PDF Pages leaves only the pages worth reading, and Compress PDF brings a heavy scan down to something you can actually attach to an email.
After the text is out
Read it against the original, then use it wherever it is going: a spreadsheet, a document, a translation request. To go the other way, note what this site cannot do — it has no way to turn text into a PDF. JPG to PDF builds a PDF from images, and for pages of text the honest options are the word processor you already have or a print-to-PDF in the browser.
Questions
How can I tell whether my PDF contains text or is a scan?
Open it and try to select a line of words. If the text highlights and can be copied, the characters are in the file and you do not need OCR. If nothing highlights and copying gives nothing, the page is a picture, so the words have to be recognised by OCR before anything can be read out of it.
Why does extracted text come out empty or as nonsense?
A PDF stores text as glyph codes, and a working tool needs the font's character map to turn those codes back into characters. Scans, some older printers and a few office suites leave that map out, so the file has text shapes but no readable text. Running OCR on the page instead is the way around it.
Can I transcribe a PDF without uploading it?
Yes, if you run the software yourself. Your browser's PDF viewer for copying, Poppler's pdftotext for a whole document, and Tesseract or OCRmyPDF for a scan all work on your own machine. Web services and phone OCR features usually send the file or the page image to a server.
How accurate is OCR, and can I use the result as it is?
Treat it as a draft that has to be checked. Recognition is guessing from shapes, so similar-looking characters, small print, furigana and vertical Japanese text are where it slips. Compare the result with the original page before you file it or send it on.