What Ia A Tagged PDF?
tagged-pdf
PDF documents can contain information about whether they are actually Tagged PDF files or not. Check out what this says about the information hidden in your PDF files.
A tagged PDF carries a hidden map of its own contents. Alongside the text and pictures you see, it records what each piece of content is - a heading, a paragraph, a list, a table cell, a caption, an image - and the order in which it should be read. The Tagged PDF field is a yes or no answer telling you whether your document has that map.
Which files carry it and where the value comes from
Only PDF documents have this field. The tags are added when the PDF is created, and whether they appear depends entirely on the program that made it and the options that were chosen.
Word processors and layout applications such as Microsoft Word, LibreOffice and Adobe InDesign can export a tagged PDF, usually behind an export option that mentions document structure or accessibility. PDFs that come from a scanner, or from a generic print to PDF driver, normally come out untagged, because at that point the content has already been flattened into text and images with no structure left to describe. Tags can also be added to an existing PDF afterwards, either by hand or with help from software, though the results vary.
What the value looks like
There are only two possible answers:
Tagged PDF: Yes- the document contains a structure mapTagged PDF: No- it does not
The field says nothing about how good the tags are. Structure produced automatically can label things wrongly, mistaking a heading for ordinary text or scrambling the order of a multi-column page, and the flag will still read Yes. A No, on the other hand, is a firm answer: there is no structure there to rely on.
It is worth reading next to the other PDF rows in the report, such as PDF Version, whether the file is linearized for fast web viewing, and any conformance marker like PDF/UA, the standard for accessible PDF documents, or PDF/A, used for long term archiving.
Why it matters
The main reason tagging exists is accessibility. Screen readers, which read documents aloud to people who cannot see them, depend on the structure map. With tags, the software knows this line is a heading, this block is a table with three columns, this picture has a description attached, and it reads the page in the right order. Without tags it is left guessing from the position of marks on the page, and an untagged PDF - especially one with columns, sidebars or tables - can come out as a jumble or as nothing useful at all.
Tagging pays off in ordinary use as well. Text copied out of a tagged PDF arrives in sensible order rather than shuffled between columns. Reflow, which lets a document rewrap to fit a phone or an e-reader instead of forcing you to pan around a fixed page, needs the structure to know what to reflow. Converting the PDF into another format keeps its headings, lists and tables instead of turning everything into flat paragraphs. And software that extracts data from documents works far more reliably when it can see where a table begins and ends.
There is one more practical angle: many public bodies and large organisations require documents they publish or accept to be accessible, and a tagged PDF is the baseline for that. Checking this field is the quickest way to know whether a document meets it. The flag itself contains no personal information whatsoever.