The problem
A field rep at a doorstep has seconds to answer an objection accurately and in the product's own words. The material exists (pricing, objections, scripts), but it lives in documents nobody can read mid-conversation. PitchPilot's copilot turns that material into something a rep can ask in plain language.
Constraints that shaped it
- Answers have to come from the product's documents, not from the model's general knowledge. A wrong price is worse than no answer.
- Each white-label brand has its own knowledge base, and one brand must never retrieve another brand's material.
- It runs on free-tier model quotas, so a global daily cap and a per-user rate limit stop one account from using up everyone's budget.
- Retrieved text is untrusted input. A document that says "ignore your instructions" has to stay data.
Architecture
The source is private, so this is the design. A smaller version of the same idea is public in rag-demo.
- IngestA Node indexer reads the brand's docs and exported objections and splits them by heading into chunks of about 3,000 characters, each with a SHA-256 content hash.
- Embed and reconcile
gemini-embedding-001with the document task type, 768 dimensions, stored in a pgvector table per tenant. Only new or changed hashes are embedded again; removed chunks are deleted. - RetrieveA Supabase Edge Function embeds the question with the query task type and calls a SQL function that runs an exact cosine search limited to the brand and returns the six closest chunks.
- AssembleChunks go into a labelled context block inside the system prompt, which tells the model to treat that block and the rep's message as data, forbids inventing prices or policies, and says what to do when nothing matches.
- GenerateGemini writes a short answer, trying three models in order of their free daily quota, lite model first, with thinking disabled.
- Respond and logThe client gets the answer plus whitelisted actions, such as opening the source document. The retrieved sources and the turn are stored for review.
- ChunkMarkdown files split into 600-character windows with 100 characters of overlap, cut at blank lines or sentence ends.
- Embed
all-MiniLM-L6-v2runs locally: 384 dimensions, normalised, no API key. - Store and retrievepgvector in Docker Compose;
match_chunks()asks for the four nearest chunks by cosine distance. - AnswerGemini receives the rules in its system instruction and only context plus question in the user turn. It cites file names or says it doesn't know from the docs.
Decisions and trade-offs
- Exact search instead of an approximate index.At a few hundred rows an exact scan is cheap and always correct, while the IVFFlat index was returning empty results. The migration records when an HNSW index becomes worth adding: tens of thousands of rows per product.Cost: this has to be revisited as the knowledge base grows.
- Reconcile by content hash instead of delete-and-reinsert.Re-inserting everything reset every timestamp, so a stale commission rate looked as fresh as a new edit. Hashing touches only what changed and lets an admin view flag pricing chunks nobody has verified in 90 days.Cost: the hash must be computed identically in JavaScript and SQL.
- A tenant registry with a foreign key instead of a hardcoded list.The old code allowed two tenants and the indexer overwrote the demo slot on every rebrand, so two white-label demos could not exist at once. The registry fixed that, and I ran an adversarial review of it for cross-tenant leaks.Cost: another table and migration to maintain.
- Thinking disabled, with a three-model fallback.The reasoning tokens of a thinking model consumed the output budget and cut answers mid-sentence. Separate free-tier quotas per model keep the copilot answering when one runs out.Cost: less deliberate reasoning, and most answers come from a lite model.
- Sources logged, not printed in the answer.A rep reading at a door does not want
[chunk 3]mid-sentence. The interface can offer a button to the source instead, and every turn records what was retrieved.Cost: the answer text carries no visible citation. - Metadata-only tracing by default.Prompts and completions contain menus, notes and customer questions. The tracing wrapper sends model, latency and token counts only; capturing content is an explicit opt-in because it would make the tracing service a data processor.Cost: debugging a bad answer needs the stored turn, not the trace.
- Local embeddings in rag-demo.No embedding key or cost, and with no Gemini key it still prints the best matching chunk, so anyone can run it.Cost: a smaller embedding model than PitchPilot uses.
What broke, and what I changed
Retrieval returned nothing
- Symptom
- The copilot kept deflecting to the FAQ for questions the documents answered.
- Cause
- An IVFFlat index built with 100 lists over roughly 250 rows, queried with one probe. Most probes landed in lists with nothing close in them.
- Fix
- Dropped the index for exact nearest-neighbour search and wrote down the size at which to add an approximate index back.
Answers stopped mid-sentence
- Symptom
- Replies were truncated even though they were short.
- Cause
- A thinking model spent the output token budget on reasoning before writing the answer.
- Fix
thinkingBudget: 0, with the reason left in a code comment.
The quota ran out during QA
- Symptom
- Requests failed partway through a run of test questions.
- Cause
- One free-tier quota shared by the demo and everything else.
- Fix
- A fallback across models that each have their own free quota, a separate key for the demo tenant, a daily cap and a per-user limit.
Grounding rules sat next to untrusted text
- Symptom
- In rag-demo the rules were concatenated into the same user turn as the retrieved chunks, so a chunk saying "ignore the above" sat beside them as an equal.
- Fix
- Rules moved to the system instruction, with tests that assert they never leak into the user turn. The change is public in the rag-demo history.
A guard flagged a correct answer
- Symptom
- While writing an evaluation dataset for output guards, a correctly grounded €99 was marked as an invented number.
- Cause
- The euro-amount pattern swallowed the full stop that ended the sentence.
- Fix
- Tightened the pattern and kept the case in the dataset. Writing the evaluation found the bug before the guard reached a live path.
Testing
- rag-demo: offline pytest in GitHub Actions for chunk coverage, normalised embeddings, a nearest-chunk sanity check and the grounding boundary. No database or API key needed.
- Output guards: deterministic checks for model JSON, stale pricing, prompt injection, refusals, cross-tenant chunks and ungrounded numbers, run against an evaluation dataset in CI. They are tested but not yet wired into the live request path.
- Isolation: Playwright tests check that the knowledge-base views cannot be read across tenants.
- Not there yet: a labelled retrieval set to measure recall, and a model-in-the-loop check of answer faithfulness. Those are the next tests I would add.
- Known gaps: in PitchPilot the retrieved chunks still sit inside the system instruction, the arrangement rag-demo moved away from; and rag-demo still creates an IVFFlat index, which at its size should be an exact scan. Both are next on the list.
Status and limits
PitchPilot's copilot sits behind the rep login, under a fictional energy-sector brand, and has not been used by a real sales team. Its source is private. rag-demo is public under the MIT licence and is the place to read the retrieval approach in code.