Skip to main content

Module metadata_quality

Module metadata_quality 

Source
Expand description

U+FFFD in resolver metadata: detect it, and repair it only from a source the user enabled (#608).

Some Crossref records carry U+FFFD REPLACEMENT CHARACTER where the publisher’s deposit lost its umlauts (Zeitschrift f\u{FFFD}r Physik, N\u{FFFD}herungsmethode). The entry still compiles, so the damage shows up only in the rendered bibliography. OpenAlex ingests Crossref and carries the same loss; Semantic Scholar, measured 2026-09-29 on the two DOIs in the issue, does not.

Repair is guarded. A candidate replaces a damaged value only when it matches it with each U+FFFD standing for one or two characters and every other character equal. So Zeitschrift f\u{FFFD}r Physik accepts Zeitschrift für Physik and refuses OpenAlex’s The European Physical Journal A (the journal’s later name), which leaves the field flagged rather than swapped for a different fact.

No new network by default. Only sources already enabled for this run (DOIGET_ENABLE_S2, DOIGET_ENABLE_OPENALEX) are asked, through the same rate limiter, allowlists and provenance log as any other call.

Structs§

QualityReport
What repair_with did.

Constants§

REPLACEMENT_CHAR
U+FFFD REPLACEMENT CHARACTER.
RESTORE_MAX_CHARS
The longest title or venue restores will align (#649 review): far past any real one, short enough that a runaway answer costs nothing.

Functions§

has_replacement_char
Whether s carries a U+FFFD.
repair
Repair m from the production Semantic Scholar and OpenAlex sources, each asked only if enabled in profile. See repair_with.
repair_with
Repair the U+FFFD fields of m that a source can supply (title and venue), asking each of sources that can serve the DOI at most once, and record each repair in [doiget].repaired_fields.
replacement_char_fields
The fields of m that carry a U+FFFD: title, authors, venue, publisher, abstract, in that order.
restores
Whether candidate is damaged with its U+FFFD characters restored: each U+FFFD in damaged stands for one or two characters of candidate (a lost UTF-8 sequence may have been one or two code points), and every other character is equal. A candidate that itself carries a U+FFFD never matches.