AI/ML · Software · 6 min read

Northstar RAG System

An assistant that would rather say “I don’t have enough evidence” than make something up.

I independently built Northstar end to end, a hands-on RAG engineering project: a retrieval-augmented assistant that grounds every answer in a nine-document corpus and refuses rather than guessing when the evidence doesn’t support an answer.

My role
I independently built the RAG system end to end as a hands-on RAG engineering project: ingestion, chunking, retrieval, grounded generation, citation validation, refusal behaviour, evaluation, testing, and Docker deployment.
Context
Independent RAG engineering project
Status
Independent build
Platform
Backend API service

Independent build

  • 9internal documents4 PDF · 2 TXT · 3 Markdown
  • 700characters per chunk100-character overlap
  • 3source formats ingested
  • Python
  • FastAPI
  • LangChain
  • sentence-transformers
  • Chroma
  • RAGAS

Roles: AI/ML Engineer · Software Engineer.

The problem

Plain LLM chat over company documents is good at sounding sure of itself. Ask it something the source material never actually covers, and it will still produce a confident, fluent paragraph, one that reads exactly like a correct answer and often isn’t. For internal tooling, that failure mode is worse than no answer at all: an employee who trusts a fabricated travel-expense rule or a fabricated escalation step has no way to know it was invented. I wanted to know what it took to build an assistant that would rather say “I don’t have enough evidence” than make something up.

Who it was for

I framed Northstar around a fictional mid-sized company, Northstar Technologies, and built its knowledge base out of nine internal documents spanning three formats on purpose: an employee handbook, an IT security policy, a remote-work policy, and a travel and expense policy as PDFs; a benefits and leave guide and a customer-support escalation procedure as plain text; and a data classification and retention standard, an engineering on-call runbook, and a platform architecture overview as Markdown. Mixing formats meant ingestion had to handle all three cleanly, not just the easy case. It also meant the documents differ the way real ones do: a PDF has pages, a Markdown file has headings, and a plain-text file has neither, so the system couldn’t assume any one structure.

My role

I built Northstar independently, end to end, as a hands-on RAG engineering project — not a course assignment or a graded exercise. I designed the corpus, the chunking and retrieval pipeline, the refusal policy, the grounded-generation prompts, the evaluation harness, and the test suite myself, and packaged the whole thing to run locally with Docker Compose.

What I built

The system extracts text from every document format, splits it into recursive chunks, embeds those chunks locally, and stores them in a persistent vector database that a FastAPI service queries at request time.

Chunking was the first real decision. I split text into 700-character chunks that overlap by 100 characters, using a recursive splitter that prefers to break at paragraph and sentence boundaries before it cuts mid-sentence. A chunk that is too large buries the one sentence that answers a question inside unrelated policy text. One that is too small loses the context that makes the sentence mean anything. The overlap is there so a rule that straddles a boundary still appears whole in at least one chunk.

The embedding model runs locally. Every chunk is embedded with a small sentence-transformers model, so no document text leaves the machine to be indexed and a question costs nothing to embed. For a corpus this size, I expected retrieval quality to depend far more on how the text is chunked than on how large the embedding model is.

The service itself is small on purpose. FastAPI and Uvicorn expose a health endpoint, so a deployment can tell whether the process is up, and a question-answering endpoint that takes a question and returns the answer or the refusal. The vector store is a persistent Chroma collection kept on a Docker volume, so restarting the stack does not mean re-embedding nine documents. Pydantic models define the request and response shapes, which keeps the contract between the API and its callers explicit and testable.

The API answers with more than text. Each response lists the sources it used, and every source carries the document name, the page, the section when the format has one, the chunk id, and the retrieval distance. A question like how many days a week someone may work remotely gets back either an answer with numbered citations or the fixed refusal, and in both cases the caller can see exactly what was retrieved.

  • Multi-format ingestion

    PDF, plain-text, and Markdown sources are all normalized into the same chunk and metadata structure before anything reaches the vector store.

  • Recursive chunking

    700-character chunks with a 100-character overlap keep enough surrounding context without losing the sentence that actually answers a question.

  • Local embeddings

    A local sentence-transformer embeds every chunk, so retrieval never depends on a per-query call to an external embeddings API.

  • Distance-filtered retrieval

    A persistent Chroma store returns candidate chunks, and only the ones under a fixed distance threshold are allowed to reach the generator.

  • Source-numbered, grounded answers

    Every generated answer cites the numbered chunks it drew from, so a reader can check the claim against the source document itself.

  • Full provenance per chunk

    Each retrieved chunk carries its document name, page, section when available, chunk id, and retrieval distance.

How it works

A query is embedded with the same local model used for ingestion, matched against the Chroma store, and only allowed to proceed to generation if at least one candidate chunk falls under a fixed distance threshold. When one does, the prompt includes the matching chunks as numbered sources so the model can cite them directly; when none does, the request never reaches the language model at all.

Document sources are ingested and chunked, embedded locally, filtered through persistent retrieval, and either grounded with citations or refused when nothing clears the distance threshold. Relationships: Document sources to Ingest and chunk; Ingest and chunk to Local embeddings; Local embeddings to Chroma retrieval; Chroma retrieval to Grounded answer (distance below threshold?); Chroma retrieval to Fixed refusal (distance below threshold?).

Decisions that mattered

  • I chose a hard distance threshold that refuses before generation runs over always calling the model and hoping the citations look plausible because an internal assistant that guesses confidently is more dangerous than one that admits it doesn't have enough evidence

  • I chose local embeddings and a persistent Chroma store over a hosted embeddings API called on every query because it keeps the retrieval path free of a recurring cloud dependency, an extra network failure point, and a per-query cost

  • I chose 700-character recursive chunks with a 100-character overlap over chunking by whole page or paragraph because smaller, overlapping chunks kept individual policies retrievable without losing the sentence that actually answers the question

Hard problems I solved

  • ProblemA fluent, wrong answer is worse than no answer at all for a question about company policy.

    FixI moved the refusal decision before generation. If nothing in the corpus clears the distance threshold, the system returns a fixed "I don't have enough evidence for that" response instead of ever calling the language model.

  • ProblemThree source formats (PDF, plain text, Markdown) each carry different structural metadata.

    FixI built format-specific extraction, including PyMuPDF for PDFs, that all converge on the same chunk record, attaching page and section metadata only where the source format actually supports it.

  • ProblemA citation is only useful if someone can verify it against the original document.

    FixEvery chunk in a response payload carries its document, page, section, chunk id, and distance, so an answer can be checked, not just trusted.

Tech stack

API & ingestion
  • Python: Service language
  • FastAPI: Query and retrieval API with Pydantic contracts
  • PyMuPDF: PDF text extraction
  • LangChain: Recursive character chunking
  • Uvicorn: ASGI server
Retrieval & generation
  • sentence-transformers: Local embedding model
  • Chroma: Persistent, distance-filtered vector store
  • OpenAI-compatible LLM client: Swappable grounded-generation backend
Evaluation & deployment
  • RAGAS: Retrieval and generation evaluation harness (dataset is roadmap work)
  • pytest: API and retrieval tests, run with httpx
  • Docker Compose: Local multi-service deployment

Outcome

Northstar works today as a functional baseline: it ingests the nine-document corpus, answers questions with numbered citations back to the source chunks, and refuses cleanly when the corpus doesn’t cover a question. I have not run a full RAGAS evaluation against it yet — building the evaluation question set is still roadmap work, and I’d rather say that plainly than imply scores exist that don’t.

What does exist is the harness. The evaluation pipeline is wired for the four RAGAS measures, faithfulness, answer relevancy, context precision, and context recall, so the work left is writing the question set, not building more tooling. The API and retrieval layers are covered by pytest tests that call the service over HTTP with httpx, and the whole stack starts with Docker Compose. The limits are just as plain. The corpus is nine documents about a fictional company, which is enough to prove the pipeline end to end but not enough to say how it would behave on a few thousand real policies.

What I learned

Grounding is a system property.

It isn’t a prompt trick or a model choice. It comes from where the refusal check sits in the pipeline, how provenance is attached to every chunk, and whether “I don’t know” is ever actually allowed to win. Get that structure right and the model choice becomes a much smaller decision.

What I’d do next

The clearest next steps are the ones I already know are missing: hybrid BM25 plus dense retrieval so exact terms aren’t only found by embedding similarity, a cross-encoder re-ranker on top of the initial retrieval pass, a PII filter that reuses the models from my ASC-PIE work, and a completed RAGAS evaluation set so “it works” becomes a measured claim instead of a description. This site’s own assistant, Ask Mohamed.AI, follows the same refuse-before-generate principle Northstar is built around — see This Portfolio for how that version works in production.