Skip to main content

paper_text

Function paper_text 

Source
pub async fn paper_text(
    base: &Url,
    id: &ArxivId,
    max_chars: Option<usize>,
    ctx: &FetchContext,
) -> Result<PaperText, FetchError>
Expand description

Fetch and extract the full text of an arXiv paper from ar5iv.

base is the ar5iv base URL (production AR5IV_DEFAULT_BASE; tests inject a wiremock origin via DOIGET_AR5IV_BASE). max_chars caps the returned section body text (char_count; the title and section headings are not counted against it) (None = no cap); truncation is flagged on PaperText::truncated, never silent. When ctx.cache_root is Some, a fresh cache entry is served from disk; otherwise the text is fetched, parsed, cached (best-effort), and one Fetch provenance row is emitted.

Never opens a PDF โ€” this is a separate fetch of the ar5iv HTML artifact (ADR-0032 D1).

ยงErrors

  • FetchError::Http for transport / status failures (a 404 / 410 collapses to NOT_FOUND at the boundary โ€” the id is genuinely absent from ar5iv).
  • FetchError::TextUnavailable when ar5iv returns a 200 with no extractable prose (char_count == 0: the paper was never converted to HTML). Distinct from NotFound: the id is valid and the PDF may be fetchable, so an agent should fetch rather than treat the ref as wrong (issue #302).
  • FetchError::SourceSchema if the ar5iv URL cannot be constructed from base + the id. (HTML parsing itself is best-effort and infallible on content โ€” see the parse_ar5iv helper.)
  • FetchError::Log if the provenance write fails (fail-closed).