PDF to XML Exact Replica Conversion
Publicada el 2026-07-17
Descripción de la oferta
I need a batch of text-based PDF files converted into XML while keeping every visual and structural detail intact. The XML will be our long-term archive copy, so page breaks, headings, tables, fonts, and any embedded images must look and read exactly as they do in the source PDFs. I can provide the original documents plus an example XSD that outlines how we currently tag pages, paragraphs, tables, and images. If you have a better schema that still produces a pixel-perfect result, I’m open to it, but fidelity comes first. Deliverables • One well-formed, validated XML file for each supplied PDF, matching the original layout and formatting. • A short read-me explaining any namespaces, attributes, or special tags used. • A reproducible workflow or script (Python, Java, PDFBox, XSLT, or similar) so we can run the same process on future documents. Acceptance criteria • Side-by-side comparison shows no missing text, mis-ordered elements, or lost styling. • Every file passes well-formedness checks and validates against the agreed XSD. Please tell me how quickly you can turn this around and which toolchain you prefer for the conversion.
Skills
Fuente original: freelancer