Tables
Preserve spanning headers, merged cells, footnotes, and captions as real table structure.
Scientific document parsing
Turn complex scientific PDFs into structured data without losing tables, equations, captions, reading order, or the relationships that carry meaning.
Whole-document understanding
Complex-layout fidelity
Structured scientific data

Built for scientific content teams
Scientific documents are different
Preserve spanning headers, merged cells, footnotes, and captions as real table structure.
Keep notation, indices, equation boundaries, and numbering intact from page to output.
Reconstruct the intended sequence across columns, floating elements, and page boundaries.
Keep figures, captions, labels, and surrounding discussion connected throughout the paper.
Reconstruct the complete paper instead of flattening one page at a time.
One connected representation of the complete paper
Understand geometry and element ownership before flattening anything into text.
Build a single representation across every page instead of parsing pages in isolation.
Represent tables, equations, figures, captions, and references as first-class objects.
Keep reading order and connections between elements intact across the entire paper.
Accuracy first
Fidelity before throughput
Whole document
Cross-page structure preserved
Connected structure
Scientific elements remain related
Accuracy-first architecture
Amanuensis combines complementary OCR and document-understanding models, then uses automated reasoning to assemble one coherent representation of the complete paper.
Send a scientific PDF through the API or a document processing workflow.
Complementary OCR, layout, and vision models read every part of the paper.
Our assembly system rebuilds structure, reading order, and cross-page relationships.
Automated reasoning checks the document for structural and semantic consistency.
Receive one coherent, machine-readable representation of the complete document.
Combine OCR, layout, and vision systems so no single model's blind spots define the result.
Reconstruct sequence across headings, columns, floating elements, and page boundaries.
Preserve headers, cells, spans, footnotes, and captions as structured data instead of flattened text.
Keep figures with captions, equations with numbering, and citations with references.
Accuracy, not just throughput
General-purpose parsers are often optimized to make large collections searchable quickly. Amanuensis is built for workflows where a misplaced equation or malformed table changes the data itself.
| Decision point | General-purpose parser | Amanuensis |
|---|---|---|
| Primary optimization | Fast, broad document ingestion | Maximum scientific fidelity |
| Document model | Pages, text blocks, and generic markdown | One connected model of the complete paper |
| Tables & equations | Often flattened, simplified, or detached | Preserved as structured scientific elements |
| Reading order | Inferred locally from individual pages | Reconstructed across columns and pages |
| Downstream result | Extracted text that still needs repair | Structured document data ready for pipelines |
Scientific content infrastructure
From publishing and repositories to life-science research, Amanuensis turns difficult papers into dependable structured inputs for the systems built on top of them.
Turn submissions and backfiles into high-fidelity inputs for production, enrichment, and discovery.
Make heterogeneous collections machine-readable without flattening the structure of each paper.
Transform literature and scientific reports into dependable data for search, knowledge systems, and analysis.
A dependable foundation
Once the paper is structured correctly, teams can power search, extraction, indexing, summarization, and future authoring workflows from a dependable foundation.
Request a pilot
Bring representative PDFs and the structure your systems need. We will define the pilot around accuracy on tables, equations, captions, reading order, and whole-document completeness.