kb-api
A service that puts one knowledge base behind two front doors — a REST API for apps and the Model Context Protocol for agents — over the same Postgres store and the same search query, so the two transports cannot drift apart.
Problem
A knowledge base is only useful to an agent if the agent can query it the way it queries everything else — as a tool, over MCP — while humans and other services still reach it over plain REST. Standing up two separate services would mean two query paths that drift apart. The goal was one search implementation, two transports, with namespace and visibility scoping enforced server side so a lower-privilege caller can never widen its own scope.
Architecture
Design decisions
- One process, one query — not two services. REST and MCP are the same FastAPI
app; the MCP surface is mounted into it and every tool calls the same
_search()function the REST route does. The rejected alternative was two deployments sharing a database — simpler to reason about per-service, but it guarantees the two query paths eventually rank results differently. Sharing one function makes divergence impossible rather than merely unlikely. - Scope enforced by caller role, request params ignored. A bot token is pinned
server side to
namespace IN ('prod','lib')andvisibility='public'; any ns/visibility passed by that caller is discarded. The rejected alternative — trusting the request and validating it — leaves the widening logic on the wire, where a bug or a crafted request can defeat it. Hard-coding the scope to the role means a lower-privilege caller has no representable way to ask for more. - Lexical ranking, no vector step — stated, not dressed up. Ranking blends a Postgres full-text score with a trigram title-similarity, and there is no embedding column. This is a deliberate scope choice for a base that is mostly keyword-searchable prose, kept honest here rather than presented as semantic search.
Numbers
Live figures, measured on 2026-09-26 against the running service and its Postgres store:
Search is hybrid lexical, not vector: the score blends a Postgres full-text rank
(ts_rank_cd over a tsvector column, parsed with
websearch_to_tsquery) with a trigram title-similarity from pg_trgm. There
is no embedding column and no semantic vector step — verified against
information_schema.columns on the live table, not assumed.
Failure modes found in production
- Auth scheme collision. The MCP mount authenticates with a bearer token at the application layer; putting HTTP basic-auth in front of it as well produced a 401 loop. Bearer had to be the only gate on that path.
- Streaming buffered by the proxy. The MCP streamable-HTTP transport needs the reverse proxy to disable response buffering and hold long-lived connections open, or SSE responses stall.
- Trailing-slash redirect. The MCP endpoint only resolves with its trailing slash; without it the proxy strips the prefix and the mount 301-redirects instead of serving.
- Dropped DB connection. The single pooled connection can be reaped by the server; the app reconnects and retries once before surfacing a 503, so a stale socket returns a result instead of a hard error.
Limits
- The ranking query is a sequential scan. Because the match clause is
fts @@ q OR similarity(title,q) > 0.1, the trigramORforces the planner to evaluate similarity on every row, so the GIN full-text index is not used for the match.EXPLAIN ANALYZEshows aSeq Scan: ~123–166 ms warm, ~622 ms cold. At 2,559 rows and 43 MB this is comfortably fast, but it scales linearly, not by index — a real ceiling, not a current problem. - A single shared async connection, not a pool: fine for the current call volume, a bottleneck under real concurrency.
- Every namespace is currently
public, so the visibility scoping is exercised only in code paths and tests, not by live private content.
What's next
- A semantic vector layer alongside the lexical query is planned; today the service is full-text plus trigram only.
- Restructure the match clause (or add a trigram index the planner can use) so the common full-text path stops falling back to a seq scan as the base grows.