Skip to main content

Module discovery

Module discovery 

Source
Expand description

External literature discovery search over OpenAlex /works?search=.

This is the front half of the #281 research loop (search → triage → expand → fetch → read → map). Unlike FsStore::search (which re-finds papers already in the local store) and unlike the citation graph walker, this module turns a free-text topic into a ranked list of candidate papers — each carrying enough metadata (title / abstract / year / venue / citation count / OA status / DOI) for an agent to triage before any PDF is fetched.

§Capability tier (ADR-0031)

Discovery search is Tier 1 OA metadata, always-on: there is no DOIGET_ENABLE_OPENALEX gate and no Cargo-feature gate. It ships in the default oa-only binary. The justification (ADR-0031 D1) is that a bounded OpenAlex query is the same network-surface risk class as the Crossref / Unpaywall calls Tier 1 already makes on every fetch: read-only OA metadata, never paywalled, never a PDF.

This is deliberately distinct from crate::sources::openalex (the #[cfg(feature = "metadata")] enrichment / referenced_works[] source used by graph, which stays Tier 2 behind DOIGET_ENABLE_OPENALEX). The Source trait is ref → FetchResult; search is query → list, so it does not fit that trait and lives here as a free function reusing only the shared HttpClient, rate limiter, and provenance log via FetchContext.

§Author / venue / publisher filters (ADR-0031 D5)

OpenAlex filters authors / sources (venues) / publishers by entity ID, not free text. So paper_search first resolves a supplied --author / --venue / --publisher name to its OpenAlex ID via a ?search= lookup against /authors, /sources, /publishers, then filters /works by authorships.author.id / primary_location.source.id / primary_location.source.publisher_lineage. The top hit is NOT taken blindly: select_entity resolves only an unambiguous name (a single hit, an exact case-insensitive name match, or a top hit that clearly out-scores the runner-up); a name matching several entities with no clear winner is a typed FetchError::Ambiguous listing the candidates, and a name matching nothing is FetchError::NotFound. The filter is never silently dropped.

§Metadata-only contract (ADR-0031 D3)

Every call here uses HttpClient::fetch_bytes (a JSON body), never fetch_pdf, and never follows an OA URL. The abstract is reconstructed from OpenAlex’s abstract_inverted_index.

Structs§

FrontierQuery
Parameters for frontier_view: the gap-spotting view that surfaces candidate papers structurally connected to a seed but not yet noticed.
FrontierResults
Results of a frontier_view query.
PaperHit
One candidate paper returned by discovery search.
PaperLinks
The cross-identifier “identity cluster” for a single work: its DOI, its arXiv preprint id (when one exists), the OpenAlex Work id, and the title.
PaperSearchQuery
A discovery-search request: the free-text query plus triage filters.
PaperSearchResults
The result of a discovery search: the hits plus the upstream total.

Enums§

DiscoverySource
Discovery backend that produced a PaperHit.
SearchSort
Ordering applied to the discovery result set.

Constants§

DEFAULT_LIMIT
Default page size when the caller does not specify --limit.
MAX_PER_PAGE
OpenAlex caps per-page at 200; requests above that are rejected by the API. build_search_url clamps to this as defense-in-depth, but the CLI rejects an out-of-range --limit up front (so the user is not silently given fewer results than asked).

Functions§

frontier_view
Surface the frontier neighbourhood of seed_doi: papers that cite the seed, ranked by age-normalized impact, ready for the agent to triage.
paper_search
Run a discovery search against OpenAlex and return ranked candidates.
resolve_links_for_doi
Resolve the PaperLinks identity cluster for a DOI via OpenAlex (/works?filter=doi:<doi>), in particular whether the work has an arXiv preprint.
zero_result_hint
Advice to attach to a paper search that matched nothing (#534).