1/** 2 * Shared MiniSearch schema for the full-text index (#276). 3 * 4 * The build step (data/lib/search-index-builder.js) and the runtime loader 5 * (frontend/App.js) must agree on idField/fields/storeFields. MiniSearch.loadJSON 6 * throws when the query-time options differ from the ones the index was built 7 * with -- the serialized index does not carry its own config -- so any drift 8 * would silently drop full-text search back to substring-only. Keeping the 9 * schema in one module makes that drift impossible instead of caught after the 10 * fact. 11 * 12 * It lives under frontend/ (not data/lib/) on purpose: frontend/ is deployed to 13 * the server, data/lib/ is build-only tooling that never ships, and the runtime 14 * side has to import this in the browser. So it stays dependency-free -- the 15 * builder and the loader each bring their own MiniSearch; this module only 16 * describes the shared shape. 17 */ 18 19// Fields MiniSearch tokenizes and ranks. author is indexed so a person named in 20// a byline (or a credited curator) is findable -- the card-only substring search 21// never reads author, which is why "Joe Amditis" returned nothing (#276). 22export const SEARCH_FIELDS = ['title', 'author', 'summary', 'concepts', 'tags', 'categories', 'body']; 23 24// The social artifact exists only to recover source text past the truncated 25// card preview. Its metadata is already covered by the in-memory substring 26// search, so indexing empty copies of those fields would add per-document 27// bookkeeping with no recall benefit (#669). 28export const SOCIAL_SEARCH_FIELDS = ['body']; 29 30// Fields stored on each hit so a result row renders without a second lookup. 31// Minimal on purpose: the full record is fetched from archive-data.json by id, 32// so storing more than the title would just bloat the index artifact. 33export const STORE_FIELDS = ['title']; 34 35export function searchIndexOptions() { 36 return { 37 idField: 'id', 38 fields: SEARCH_FIELDS, 39 storeFields: STORE_FIELDS, 40 }; 41} 42 43export function socialSearchIndexOptions() { 44 return { 45 idField: 'id', 46 fields: SOCIAL_SEARCH_FIELDS, 47 storeFields: [], 48 }; 49}
Line numbers count LF bytes from the start of the resource, as the search results do. Vendor segments are library code the classifier recognised; they are stored but not indexed. Bytes are shown as Latin1 characters, one per byte.