PDF transcription: getting the text out of a PDF

"Transcribing a PDF" covers two very different jobs. Sometimes the document already contains its text and you simply cannot see it — then you are copying. Sometimes the PDF is a scan, a picture of a page, and the characters have to be recognised first — then you need OCR. Nearly every web service offering either one sends your file to a server, which is why this site will not offer one; it also has no text extractor and no OCR engine, so there is no button here that does it. Here is what actually works on your own machine.

One word of warning before the rest: in Japanese 「文字起こし」 usually means transcribing a recording. This site has no audio or video tools of any kind, so if you want a meeting or a video turned into text, none of the below is for you — the guide is about a PDF you are already holding.

First: does this PDF contain text, or a picture of text?

Open it in your browser (drag the file onto the window) and try to select a line of words with the mouse.

A few things can make real text look absent: text drawn as curves or outlines, a scanned page that was then compressed flat, or a file that only opens in a specific reader. If the file opens as a blank page, it is also worth checking the viewer's page count — a one-page document that renders as empty is usually a scan seen from the wrong side, not a broken file.

If the text is there: copy it out

One or two pages, nothing to install.

  1. Open the PDF in Chrome, Edge, Safari or Firefox.
  2. Select everything with Ctrl+A (Cmd+A on macOS), or drag across the part you need.
  3. Copy with Ctrl+C and paste into a text editor, Word or Pages. Pasting as plain text drops the formatting and leaves plain lines of text; columns come out left to right, which is not the reading order for vertical Japanese.

A whole document: pdftotext from Poppler. It is free, it runs on your own machine, and the document is never uploaded:

One thing worth knowing: a PDF stores text as codes for glyphs, not as readable characters, and the tool needs the font's character map to turn them back into words. Files produced by scanners, by some older printers and by a few office suites are missing that map, and then the extracted text comes out empty or as nonsense. That is a property of the file, not a bug in the tool.

If it is a scan: OCR is the missing step

OCR (optical character recognition) reads the picture of a word and decides which character it is. It is the only way to get words out of a scan, and it is a separate program rather than something a browser does for free.

Always read the result. OCR is guessing from shapes, so it is a draft, never a transcript to file without checking. The guesses that go wrong most often:

Uploading is the part to think about

It is worth being clear about where the work happens. Browser OCR, phone camera text recognition and the web services that offer "PDF to text" send the document or the page image to a server to be processed. Local tools — your own viewer, pdftotext, Tesseract, OCRmyPDF — keep the file on your machine, which matters when the document is a contract, a medical record, a payslip or anything with somebody else's personal details in it.

Getting a scan ready so OCR reads it well

This is where the tools on this site do help, because they all work on your own machine:

After the text is out

Read it against the original, then use it wherever it is going: a spreadsheet, a document, a translation request. To go the other way, note what this site cannot do — it has no way to turn text into a PDF. JPG to PDF builds a PDF from images, and for pages of text the honest options are the word processor you already have or a print-to-PDF in the browser.

Questions

How can I tell whether my PDF contains text or is a scan?

Open it and try to select a line of words. If the text highlights and can be copied, the characters are in the file and you do not need OCR. If nothing highlights and copying gives nothing, the page is a picture, so the words have to be recognised by OCR before anything can be read out of it.

Why does extracted text come out empty or as nonsense?

A PDF stores text as glyph codes, and a working tool needs the font's character map to turn those codes back into characters. Scans, some older printers and a few office suites leave that map out, so the file has text shapes but no readable text. Running OCR on the page instead is the way around it.

Can I transcribe a PDF without uploading it?

Yes, if you run the software yourself. Your browser's PDF viewer for copying, Poppler's pdftotext for a whole document, and Tesseract or OCRmyPDF for a scan all work on your own machine. Web services and phone OCR features usually send the file or the page image to a server.

How accurate is OCR, and can I use the result as it is?

Treat it as a draft that has to be checked. Recognition is guessing from shapes, so similar-looking characters, small print, furigana and vertical Japanese text are where it slips. Compare the result with the original page before you file it or send it on.