Skip to main content

Module paper_text

Module paper_text 

Source
Expand description

Full-text extraction of an arXiv paper from ar5iv’s LaTeXML XHTML (the #281 “read” step; ADR-0032).

This is the step that lets an agent actually read a paper without an external pdf-to-text tool. It is deliberately distinct from PDF content processing (permanent non-goal #1 / ADR-0003): the PDF blob is never opened. Instead doiget fetches a separate, already-structured artifact — the publisher-rendered HTML — and extracts text from it. ADR-0032 D1 records the boundary: PDF-blob parsing and OCR stay permanently out of scope; structured HTML/XML full text is in scope.

§Source (ADR-0032 D3)

PR4 ships one source, ar5iv (ar5iv.labs.arxiv.org/html/<id>), which renders arXiv papers as LaTeXML XHTML. The fetch goes through the dedicated "ar5iv" source key (see crate::http::fulltext_allowlist) so the provenance trail distinguishes ar5iv full text from the arXiv PDF/Atom API. PMC / Europe PMC JATS is a planned follow-up source.

§Capability tier (ADR-0032 D2)

Tier 1 OA metadata, always-on: no env gate, no Cargo feature gate. Read-only, open-access, never a PDF reinterpretation — same posture class as discovery search (ADR-0031). Ships in the default oa-only binary.

§Caching (ADR-0032 D4)

Extracted text is cached at <cache_root>/text/<safekey>.json (the doiget-private cache root, docs/CACHE.md) — not the shared ~/papers/ store (docs/STORE.md), so no cross-tool coordination is needed. The cache holds the full text; max_chars truncation is a view applied on return, so one cached entry serves any max_chars. Best-effort: a miss / parse error / write failure degrades to a re-fetch, never an error (mirrors crate::resolver_cache).

§Extraction

A quick-xml walk (the same parser the arXiv Atom path uses) splits the document into { heading, text } sections on h1–h6, skips script / style / math subtrees — capturing each <math>’s alttext (the LaTeX source) as inline \(…\) text so formulae read cleanly rather than as MathML noise — and normalizes whitespace. Extraction is best-effort: it supplies the text it can and flags truncation; it does not promise faithful reconstruction.

Structs§

PaperText
Extracted full text of an arXiv paper.
TextSection
One { heading, text } section of an extracted paper.

Enums§

TextSource
Which structured full-text source produced a PaperText.

Constants§

AR5IV_DEFAULT_BASE
Production ar5iv base. Overridable via DOIGET_AR5IV_BASE (test wiremock origin), mirroring the DOIGET_ARXIV_BASE override.

Functions§

paper_text
Fetch and extract the full text of an arXiv paper from ar5iv.