Drop a PDF and convert its text layer into a Markdown file and a DOCX file, entirely in your browser. Reading order is recovered geometrically with a band-first recursive XY-cut: each page is split into full-width horizontal bands first, table rows are detected before any column split so aligned tables stay intact, and what remains is cut at column gutters only where real columns are flanked on both sides, so two-column text reads sequentially instead of splicing left and right fragments into each line. Headings come from font-size tiers plus all-caps section headings, paragraphs rejoin end-of-line hyphens, and bulleted and numbered lists are detected as before. A page with no extractable text layer, a scanned or photographed page, is detected and reported inline instead of producing silent empty output; optical character recognition is out of scope. When a text run reaches the page's right boundary it can be cut off mid-word by the text layer itself (a pdf.js extraction behavior, still present upstream); such pages are flagged inline in both outputs rather than shipping the clip silently. The pdf.js extraction engine is pinned and inlined workerless, making this the heaviest page on the site (about 2 MB) – acknowledged here rather than engineered away.
@licstart/@licend banner, unmodified, at the top of each inlined block). Everything past extraction, line grouping, heading/list/table detection, Markdown rendering, the OOXML template, and the ZIP writer, is this tool's own code, not vendored. See credits.html for the suite-wide attribution ledger.