pub async fn paper_text(
base: &Url,
id: &ArxivId,
max_chars: Option<usize>,
ctx: &FetchContext,
) -> Result<PaperText, FetchError>Expand description
Fetch and extract the full text of an arXiv paper from ar5iv.
base is the ar5iv base URL (production AR5IV_DEFAULT_BASE; tests
inject a wiremock origin via DOIGET_AR5IV_BASE). max_chars caps the
returned section body text (char_count; the title and section
headings are not counted against it) (None = no cap); truncation is
flagged on
PaperText::truncated, never silent. When ctx.cache_root is Some,
a fresh cache entry is served from disk; otherwise the text is fetched,
parsed, cached (best-effort), and one Fetch provenance row is emitted.
Never opens a PDF โ this is a separate fetch of the ar5iv HTML artifact (ADR-0032 D1).
ยงErrors
FetchError::Httpfor transport / status failures (a 404 / 410 collapses toNOT_FOUNDat the boundary โ the id is genuinely absent from ar5iv).FetchError::TextUnavailablewhen ar5iv returns a 200 with no extractable prose (char_count == 0: the paper was never converted to HTML). Distinct fromNotFound: the id is valid and the PDF may be fetchable, so an agent should fetch rather than treat the ref as wrong (issue #302).FetchError::SourceSchemaif the ar5iv URL cannot be constructed frombase+ the id. (HTML parsing itself is best-effort and infallible on content โ see theparse_ar5ivhelper.)FetchError::Logif the provenance write fails (fail-closed).