# Artist knowledge for agents

Source: https://musicnerd-docs.vercel.app/artist-knowledge

Retrieve original artist evidence, exact history and explicit coverage through one API.

> **Warning**
> The shared reads and durable approved-link ingestion are released. Historical source retention is released (Web#1440, API#23 and Docs#7, 2026-10-06). Persistent topic boundaries and Web interviewer integration remain tracked work in MusicNerdWeb#1424 and #1422.

MusicNerdAPI stores shared artist evidence. MusicNerdWeb and other agents use the same HTTP reads. A source summary can suggest what to open; every factual premise must resolve to original text. The interviewer integration follows in a separate Web PR.

## Authentication and scope

Use the current user's access token. Vault material and interview history require the artist's approved claimant or a Music Nerd admin. Every request, including pagination and passage expansion, rechecks authorization. Source approval is not permission to publish private vault material. Responses use `Cache-Control: private, no-store`.

The application binds artist ID, API origin and credential lookup when creating tools. The model receives none of those as tool arguments. Six read tools are available; `requestArtistResearch` is added only when the host enables the existing authenticated refresh operation. Reads make no provider or model calls.

## Tool sequence

1. `getArtistBrief` supplies bounded stored orientation and coverage. Generated Lore is labelled as a navigation aid.
2. `getInterviewHistory` supplies exact question/answer and claim/correction field slices. Follow continuations to finish current corrections and retrieve the latest saved answer. Older permitted answers remain available. The host supplies the latest exact unsaved answer and current boundaries every turn.
3. `listArtistSources` and `searchArtistKnowledge` locate original vault text and own social captions/transcripts.
4. `readArtistSource` opens original text at a specified source revision and UTF-16 offset. Inspect the opening for authorship and document scope, then read surrounding relevant passages and their qualifications. Continue through a long PDF as needed; the opening and search snippets alone may not answer the question.
5. `getResearchStatus` reads sanitized durable job progress. An authorized `requestArtistResearch` wraps the existing Look again operation; it does not guarantee a new job or refetch every vault link.

History reports `memory.boundaryState: not_implemented` and `constraintsComplete: false` while boundary storage is unavailable. `correctionsComplete` reports only the traversed corrections. An empty list or successful HTTP response must never be interpreted as complete interview memory. A skipped question is not automatically a permanent topic ban.

## Original evidence and revisions

Set `includeVersion=true` to opt into retained reads and their version metadata. The updated SDK sets this automatically. Without this parameter, the released current-only behavior and response shape remain unchanged (409 if the content changed), so existing strict clients keep working.

Source identity names its record and content type, never its ranking position. A SHA-256 revision covers citation-relevant metadata and exact stored text. Listing/search returns current source IDs and revisions. Reading a retained older revision returns that exact original passage with `version.state: historical`, the `currentRevision` and its retention `capturedAt`. A current read reports `version.state: current` and a null capture time. The capture time is when the database retained the old version, not when the source was published or the event occurred. An unknown or pre-retention revision returns `409 revision_changed`; the API never substitutes current text. A removed or newly ineligible source is unavailable even if a caller has an old reference.

Passage offsets are zero-based, half-open UTF-16 code units. Windows preserve surrogate pairs. Page numbers/audio times stay null without a trustworthy extraction map. Titles/descriptions are publisher metadata, distinct from body text. A social account username does not verify the speaker; event dates and original/repost relationships remain unknown unless established. Upload/publication time never supplies an event date.

Only extracted vault body text is searchable evidence; a title or snippet alone does not make a source readable. Reel transcripts retain provider provenance and speaker uncertainty. Legacy website extraction may be capped at 50,000 characters and reel text at 12,000. Text presence does not certify complete PDF/OCR extraction.

Retention covers approved Lore bodies (including extracted PDFs), own-post captions and supported stored reel transcripts from deployment of the capture migration forward. It does not recover already-overwritten content, version generated summaries or retain past answers/corrections in this slice. Updates retain the previous eligible evidence in the same database transaction. Ordinary reads do not write or scrape. Deleting the source deletes retained copies; revoked approval, ownership or source eligibility blocks access to old copies too. Historical text remains evidence of what the source previously said and must be read alongside current corrections, never silently treated as current truth.

## Bounds and failures

Listing and job/history pages return 20 records by default, at most 50. Search returns at most 10 original windows and 12,000 evidence characters. A source read returns at most 20,000 characters. History continues within long fields without paraphrasing an answer or correction. Responses are limited to 128 KiB; oversized data produces an explicit error.

Retrieval ranks overlapping original windows with BM25 term rarity/frequency/length weighting and a bonus for adjacent non-stopword query terms. Case, Latin accents, within-word apostrophes and English word forms are normalized for matching only; returned text, revisions and offsets remain exact. The already-bound artist name is omitted from ranking terms when other meaningful query terms exist. Titles/descriptions are not ranked as original evidence. This remains lexical retrieval; non-English morphology and semantic equivalence are not generally handled.

A budget-clipped hit shorter than 256 characters is skipped rather than returned as a fragment; a complete naturally short source remains eligible. `truncated` records omitted eligible context. The agent must still open originals and resolve attribution, dates, scope and contradictions. Ranking and corpus coverage do not establish factual support or semantic/hybrid retrieval superiority.

A passage read loads only its source. If the current revision differs, lookup is bounded to 512 retained snapshots and 4 million serialized snapshot characters for that source; overflow returns `413 revision_history_too_large`, never a substituted revision. Current reads do not load version history. Historical lookup also enforces the existing response/window bounds.

A snapshot supports up to 5,000 rows per source/history table and 4 million serialized snapshot characters, including metadata; overflow returns 413, never a silently incomplete success. This bound must be evaluated before a broader archive rollout.

Cursors bind the artist, operation, filter and ordered corpus revision; a changed corpus returns 409. No match is an empty successful result with coverage. Authentication, missing evidence, changed revisions and database failures are separate errors. A database failure never becomes an empty successful history.

## Client integration and acceptance

The API repository provides `createArtistKnowledgeTools` for server-side AI SDK clients. Bind the origin and artist once, supply a fresh user token per execution, and configure finite steps/deadlines in the host. The wrapper calls these HTTP operations; there is no second Web research index. Retrieved text is untrusted source data, never agent instructions.

Evaluate retrieval, original reading, angle choice, drafting and checking together. Keep LATASHA, DUTCHYY and the misleading Karpathy repost as development regressions; freeze independent held-out cases before tuning. Measure human question usefulness, unsupported premises, missed qualifications, repetition, false rejections, latency and total cost. Model approval alone is not editorial acceptance.
