Accessible PDF documents: creating them and checking them

Statutes, minutes, application forms, official gazettes, annual reports – much of it sits online as a PDF, and most of it is unusable for blind people. This article covers who is obliged to act, what makes a PDF accessible, how that is decided in the source document, what you check the result with, and what to do about the hundreds of legacy files whose source document is long gone.

Published on 1 October 2026 · Libration editorial team · 12 min read

First, the question almost nobody asks: does this have to be a PDF?

A PDF is the image of a printed page. Fixed width, fixed line breaks, fixed type size – that is its purpose, and that is exactly what stands in the way of accessibility. Double the text size on a web page and you get a new line break; do it in a PDF and you get a magnifying glass, with which you push every single line to the right and back again.

So the first piece of advice is almost always: if the content is meant to be read, publish it as a web page. A set of statutes, a schedule of fees, minutes, a press release – all of that is more accessible as HTML, easier to update, easier to find, and usable on a phone in the first place.

A PDF is the right choice when one of these three things applies:

  • The document has to be archived unchanged – annual accounts, a certificate, a formal decision.
  • The document is filled in and sent back, and there is no online form for it.
  • The layout carries meaning that HTML cannot reproduce: maps, plans, sheet music.

And here is where most people take this thought one step too few: an accessible HTML version alongside does not replace an inaccessible PDF. Leaving the document online and adding a web page with the same content does not discharge the obligation – it just creates two versions. Moving to HTML helps precisely when the PDF disappears in the process.

A production reason, not a reader's need Many organisations publish PDFs because the file existed for print anyway. The opening question is therefore not “how do I make this PDF accessible”, but “why is this a PDF”.

Who is obliged to act

In the EU there are two separate legal tracks. They cover different groups, and they have two different cut-off dates. That is the most common confusion of all.

Public sector bodies

Public sector bodies are covered by the Web Accessibility Directive, Directive (EU) 2016/2102, as transposed into national law. The technical reference is the harmonised standard EN 301 549, whose clause 10 deals explicitly with “non-web documents” – PDFs, office files, everything that is downloaded rather than viewed.

That PDFs are meant is not an interpretation, it is in the text. The directive defines “office file formats” as content that is “not primarily intended for use on the web, that are included in web pages, such as Adobe Portable Document Format (PDF), Microsoft Office documents or their (open source) equivalents”.

Exempt are “office file formats published before 23 September 2018, unless such content is needed for active administrative processes relating to the tasks performed by the public sector body concerned”. So the legacy stock may stay – but not if somebody still needs it today in order to apply for something. A building permit form from 2016 that is still being filled in and submitted is not exempt. Minutes from 2016 are.

Businesses under the European Accessibility Act

The European Accessibility Act – Directive (EU) 2019/882 – does not apply to every business, but to specific products and services: e-commerce, banking, telecommunications, passenger transport and e-books among them. Who is covered and who is not is the subject of our article on the European Accessibility Act.

A different cut-off date applies here: the Act has applied since 28 June 2025. National transpositions typically exempt office documents published before that date – in Germany, for instance, section 1(4) of the BFSG does exactly that – so check the wording in your own country. For everything published afterwards, information about the service has to be available in formats “suitable for generating alternative assistive formats”. A PDF that consists only of images does not meet that.

One point that is routinely overlooked: the exemption attaches to the act of publishing, not to the file. Swap an old document, re-upload it or update its content and you have published it anew – and lost the exemption. When clearing up a legacy stock that is the single most important consideration, and we come back to it below.

Associations, clubs and everybody else

A small association without an online shop is usually covered by neither. That does not make it irrelevant: a membership form that a blind member cannot fill in is a problem with or without a law, and funders increasingly ask. If you are in this group, start with the documents that are actually used: the application form, the registration, the statutes.

What makes a PDF accessible

An accessible PDF differs from an ordinary one not in how it looks but in a second, invisible layer: the tag tree. It describes what the things on the page mean – this is a second-level heading, this is a table cell with a column header, this is a figure with this description. Without that layer, a screen reader sees only letters lying somewhere on a surface.

Building blockWhat it means
TagsEvery piece of content is marked up: heading, paragraph, list, table, figure.
Reading orderThe order in the tag tree, not the arrangement on the paper. In multi-column layouts the two almost always differ.
Heading levelsH1, H2, H3 without gaps. They are the table of contents blind readers use to move through a document.
Document titleIn the metadata, not the file name. It appears in the window title and is the first thing read out.
LanguageFor the document as a whole, and separately for passages in another language. Otherwise English is read with a German accent, or the other way round.
Alternative textFor every figure that carries information; purely decorative items are marked as “artifacts” and skipped. The rules are in our article on alt text.
TablesWith real header cells, and for complex tables with cells associated to their headers. Tables that only build the layout are not tables.
LinksWith meaningful text. “More information” and a bare URL are equally useless.
Form fieldsWith labels and tooltips, in a sensible tab order, required fields not marked by colour alone.
BookmarksFrom about ten pages up. In long documents the single most important way to navigate.
Contrast and typeThe same values as on the web: 4.5 : 1 for body text. And real text instead of text in an image.

PDF/UA: the standard behind it

What the WCAG are on the web, PDF/UA (“Universal Accessibility”) is for PDF:

  • PDF/UA-1 – ISO 14289-1, published in 2012 and revised in 2014. This is the version public bodies, tenders and testing tools still refer to.
  • PDF/UA-2 – ISO 14289-2, published in 2024 and built on PDF 2.0. New additions include MathML for formulae, dedicated structure elements for footnotes and asides, and more modern Unicode handling. In practice PDF/UA-1 is still the norm.

Alongside it sits the Matterhorn Protocol from the PDF Association, which translates the standard into 31 checkpoints with 136 failure conditions. It is the basis of nearly every testing tool – and at the same time the best evidence that testing software alone is not enough: 45 of those 136 conditions cannot be decided by a machine at all. Whether an alternative text exists is something software can tell you. Whether it is right is not.

PDF/UA and WCAG are not alternatives The WCAG apply to PDFs as well – the W3C maintains its own PDF techniques for them. PDF/UA is the technically more precise statement of the same requirements for this one file format.

Accessibility is decided in the source document, not in the PDF

The most common and most expensive mistake: export first, then “make it accessible”. Tagging after the fact is manual work, and it has to be repeated after every content change. Almost everything that matters is decided before.

Microsoft Word

  1. Use styles, do not format. A heading is “Heading 1” – not bold, larger and centred. Only the style becomes a tag.
  2. Alternative text on every figure; mark decorative images as decorative.
  3. Tables: mark the header row using the table tools. No merged cells, no layout tables.
  4. Document title and language in the document properties.
  5. Export via File → Export → PDF/XPS and enable Document structure tags for accessibility under Options.

What you must not do: pick a PDF printer under Print. That discards the entire structure. The result looks exactly like the original and is empty inside.

LibreOffice

File → Export As → Export as PDF, General tab, option Universal Accessibility (PDF/UA). LibreOffice also checks for typical problems during export and reports them. For plain text documents this is a surprisingly good route.

Adobe InDesign

The most laborious but most reliable route for designed documents. Three things matter: map paragraph and character styles consistently onto tags, set the reading order in the Articles panel rather than by position on the page, and enable tag creation on export.

Scanned documents

A scan is an image. Without text recognition it contains not a single letter a screen reader could read out – it is not “poorly accessible”, it is not accessible at all. OCR is the first step, the result has to be proofread, and only then does the structural work begin. Where the source file still exists, re-exporting is almost always faster than rescuing the scan.

Checking: what with, and what the check does not tell you

ToolWhat it is
PAC (PDF Accessibility Checker)The de facto standard in German-speaking countries, from axes4. Free and without registration, but Windows only. Checks against PDF/UA and WCAG, shows the tag tree and a preview of what a screen reader would read out.
veraPDFOpen source, free, cross-platform, validates PDF/A and PDF/UA (parts 1 and 2). The choice when there is no Windows in the building, or when the check has to run automatically.
Adobe Acrobat ProCommercial, with its own accessibility check and the tools to correct tags by hand. For fixing a finished PDF it is effectively indispensable.
A screen readerThe only check that really counts. NVDA is free. Ten minutes of listening say more than any report.

And the sentence that appears in no test report: “no errors found” does not mean “accessible”. A tool establishes that every figure has an alternative text. Whether any of them describes what is actually there, it does not establish. The same goes for heading levels, reading order and link text: formally correct and substantively useless are indistinguishable to a checker.

And the legacy stock? Remediate rather than rebuild

Everything so far applies to documents you create today. The real problem is a different one: a website that has grown over years holds two, sometimes three hundred PDFs. The source files are half lost, the authors have left, and nobody knows which of these files anyone still opens. This is where most initiatives fail – not on the technology, on the volume.

That this is not an isolated case is shown by Allyant's PDF Accessibility Index: of roughly 645,000 PDF files checked across more than 770 websites, 94.75 per cent were not accessible (2025–2026 benchmark report, published in March 2026). This is not the failure of a careless few. It is the normal state of affairs.

That is what we built Libration Docs for – deliberately in two forms, because far from every website runs on WordPress:

  • As a web application on libration.io. You upload a document, get the report and can have it remediated. Nothing beyond a browser is needed – no WordPress, no plugin, no installation. That holds for any website, whether it runs on TYPO3, Joomla, a shop system or a bespoke build, and equally for documents that are not online at all yet.
  • As a Documents tab in the Libration Accessibility plugin. If you use WordPress, the next release brings the same function directly on the media library: every PDF is managed as part of a stock, with a status per file, instead of being uploaded one at a time.

The check behind both is the same one.

Checking is free. veraPDF validates against the PDF/UA-1 profile, and a layout analysis is added to it: where are structure tags missing, where is the reading order unclear, which figures have no alternative text, is there a text layer at all. Every finding names the place and the step required. That answers the question every clean-up starts with: how bad is it, and where do I begin?

Remediation rebuilds the structure – tags, reading order, headings, lists, tables with header cells, document language, document and window title, artifacts for headers and footers. For figures without alternative text the AI makes a suggestion, explicitly as a suggestion. veraPDF then runs a second time so that the difference the pass made is there in black and white. You pay only for what is remediated; checking stays free.

And before all that there is a trial run without an account: upload one of your own PDFs, have it checked, have part of it remediated, look at the report and the resulting file. At most half the pages are remediated and never more than ten, three runs per day. The fastest way to judge all this is your own document, not our description of it.

The traffic light is the important part – and its middle state is the important part of that. A document is not checked, machine-processed, unverified or checked and released. After remediation it deliberately does not go green. A machine cannot judge whether the reading order makes sense, whether a suggested alternative text says the right thing, or whether a complex table remains comprehensible. Only once a person has reviewed and released it does the document count. That is inconvenient, and it is the reason we trust the result: a tool that turns everything green after a single pass is selling you a feeling, not a state.

Replace or place alongside – and here the exemption from further up comes back into play. By default a new file with the suffix -accessible is created next to the original; existing links keep working. Replacing the original is tidier, but the file has then been published anew, and the exemption for legacy documents falls away. That decision should be taken deliberately, not in passing – and it arises in exactly the same way if you use the web application and upload the resulting file yourself.

Data protection Processing runs on a server in Germany. The source file is deleted immediately after processing, the result after 14 days. For alternative texts only the individual figure is sent to the AI model, not the whole document.

And the limit, so that it does not surface only in production: remediation can add structure that can be derived from the document. It cannot turn a scan without OCR into a readable document, it cannot detect a reading order that is wrong in substance, and it can propose an alternative text but not vouch for it. For the documents that really matter – forms, applications, anything somebody has to fill in – a human pair of eyes remains mandatory.

The mistakes that come up most often in practice

  • The scan without text recognition, especially statutes and older decisions.
  • No tag tree at all, because the file was exported through a PDF printer.
  • Reading order by layout. In two-column documents one line from each column is then read out alternately.
  • Tables without header cells – and layout tables marked up as tables.
  • Missing language setting, with the result that the text is read out with the wrong pronunciation.
  • Forms without field labels. Of all places, where it hurts most.
  • Text as an image, typically scanned signatures, logos with a claim, and infographics.
  • The file name as the document title, so “annex_3_final_v2.pdf” is the first thing you hear.

An order of work that has proven itself

  1. Take stock. List every PDF, with its date and its page views.
  2. Weed it out. Anything out of date goes. The cheapest progress there is.
  3. Convert. Anything that is read rather than filled in becomes a web page – and the PDF disappears in the process, otherwise it achieves nothing.
  4. Prioritise. Of what is left, forms and applications first. That is where it is decided whether somebody can take part at all.
  5. Re-export where the source exists. Remediate where it does not.
  6. Check and listen. Testing tool, then screen reader, then release.
  7. Write the rule down. Anybody publishing a PDF in future follows a one-page guide. Without that step the backlog starts over the day after the clean-up.
Know where you stand in two minutes Upload a document at libration.io/en/docs/ – no account, no WordPress, no credit needed. You get the report and see on part of the pages what remediation makes of it. And if you want to know how the site around it is doing, the website scan is free as well.

Sources and notes

  • Directive (EU) 2016/2102, Article 1(4) and Article 3(26) – EUR-Lex, CELEX 32016L2102
  • Directive (EU) 2019/882 (European Accessibility Act), applicable since 28 June 2025; German transposition: BFSG, section 1(4)
  • EN 301 549, clause 10 “Non-web documents”. The legally relevant version is the one cited in the Official Journal of the EU; most recently that was V3.2.1 (March 2021, with WCAG 2.1). ETSI published V4.1.1 in September 2026, which adopts WCAG 2.2.
  • ISO 14289-1 (PDF/UA-1, 2012/2014) and ISO 14289-2 (PDF/UA-2, 2024)
  • Matterhorn Protocol 1.1, PDF Association – 31 checkpoints, 136 failure conditions
  • Allyant, PDF Accessibility Index, surveyed March 2026
  • PDF Accessibility Checker (PAC), axes4 – free, Windows
  • veraPDF – open source, PDF/A and PDF/UA parts 1 and 2

This article describes the legal situation and is not legal advice.