Expand description
Full-text extraction of an arXiv paper from ar5iv’s LaTeXML XHTML (the #281 “read” step; ADR-0032).
This is the step that lets an agent actually read a paper without an external pdf-to-text tool. It is deliberately distinct from PDF content processing (permanent non-goal #1 / ADR-0003): the PDF blob is never opened. Instead doiget fetches a separate, already-structured artifact — the publisher-rendered HTML — and extracts text from it. ADR-0032 D1 records the boundary: PDF-blob parsing and OCR stay permanently out of scope; structured HTML/XML full text is in scope.
§Source (ADR-0032 D3)
PR4 ships one source, ar5iv (ar5iv.labs.arxiv.org/html/<id>),
which renders arXiv papers as LaTeXML XHTML. The fetch goes through the
dedicated "ar5iv" source key (see crate::http::fulltext_allowlist)
so the provenance trail distinguishes ar5iv full text from the arXiv
PDF/Atom API. PMC / Europe PMC JATS is a planned follow-up source.
§Capability tier (ADR-0032 D2)
Tier 1 OA metadata, always-on: no env gate, no Cargo feature gate.
Read-only, open-access, never a PDF reinterpretation — same posture
class as discovery search (ADR-0031). Ships in the default oa-only
binary.
§Caching (ADR-0032 D4)
Extracted text is cached at <cache_root>/text/<safekey>.json (the
doiget-private cache root, docs/CACHE.md) — not the shared
~/papers/ store (docs/STORE.md), so no cross-tool coordination is
needed. The cache holds the full text; max_chars truncation is a
view applied on return, so one cached entry serves any max_chars.
Best-effort: a miss / parse error / write failure degrades to a
re-fetch, never an error (mirrors crate::resolver_cache).
§Extraction
A quick-xml walk (the same parser the arXiv Atom path uses) splits
the document into { heading, text } sections on h1–h6, skips
script / style / math subtrees — capturing each <math>’s
alttext (the LaTeX source) as inline \(…\) text so formulae read
cleanly rather than as MathML noise — and normalizes whitespace.
Extraction is best-effort: it supplies the text it can and flags
truncation; it does not promise faithful reconstruction.
Structs§
- Paper
Text - Extracted full text of an arXiv paper.
- Text
Section - One
{ heading, text }section of an extracted paper.
Enums§
- Text
Source - Which structured full-text source produced a
PaperText.
Constants§
- AR5IV_
DEFAULT_ BASE - Production ar5iv base. Overridable via
DOIGET_AR5IV_BASE(test wiremock origin), mirroring theDOIGET_ARXIV_BASEoverride.
Functions§
- paper_
text - Fetch and extract the full text of an arXiv paper from ar5iv.