Your senior staff answer the same 30 questions. Every single week.
James Whitfield, a senior partner at Thornfield Advisory, fielded the same questions every week: where's the onboarding checklist, what's the standard fee structure, which engagement template should I use. Eight years of institutional knowledge existed across Confluence, Google Drive, and an S3 archive, but finding any of it meant interrupting someone who already knew. AskBase connects all three, maps who can see what, and answers in Slack with a design target of under 10 seconds, cited, permission-aware, auditable. By day ninety of the scenario, gap detection had surfaced 14 processes that were documented nowhere.
COMPOSITE_DEPLOYMENTS Composite scenarios: fictional companies, built on our platform, with outcomes stated as design targets.
BUILT_ON AskBase · available now
When Knowledge Lives Inside People
Thornfield Advisory had grown from 40 people to 165 in six years. The knowledge that built the firm (how engagements were scoped, how clients were onboarded, how deal memos were structured) was documented in bursts, by whoever had time, in whatever tool was current. By 2024, institutional knowledge was scattered across three Confluence spaces, two Google Drive folders, an S3 bucket of archived files left over from an acquisition, and the heads of seven senior partners who had been there from the beginning.
The Senior Partner
- Fourteen Slack messages before 10am. Twelve of them are questions a junior analyst could answer, if they knew where to look. The answers exist in documents nobody can find.
- "What's our standard NDA turnaround?" is asked seventeen times a month. The answer is in a Confluence page from 2019. Nobody knows it exists.
- Three hours per week answering questions that are not in anyone's job description to answer, including mine. That is 150 hours per year of senior partner time spent being a search engine.
The New Analyst
- Week one: asked three questions before realising the answer was "ask Sarah." Week three: realised Sarah doesn't always know either. She asks David. Week six: stopped asking and started guessing.
- The onboarding documentation says "see Confluence." Confluence has 847 pages across three spaces. Searching returns 40 results sorted by last-modified date with no signal for which is current.
- Made a client onboarding error in month two. The correct procedure was documented, in a page nested four levels deep under a space created for a project that ended in 2021.
Six Stages From Documents to Intelligence
The system does not replace documentation. It makes existing documentation findable, citable, and queryable, by anyone with the right permissions. The design target for AI cost, with 165 users asking approximately 400 questions per month: under $0.80.
The Knowledge Pipeline
Connect
The system connects to Confluence, Google Drive and S3 through their APIs and ingests pages, PDFs, Word documents, Google Docs and Sheets, and plain text. New and updated documents sync automatically, and a full resync removes documents deleted at the source.
Map Permissions
Before a single chunk is embedded, each document is given permission groups from its source: a Confluence page takes the space it lives in, and a Google Drive file the groups and domain it is shared with. An admin maps those labels to the firm's groups, so a page in the Deal Team Alpha space is visible to Deal Team Alpha only, and a Drive folder shared with everyone is available to All Staff. A document that arrives with no label goes to All Staff; today that includes a file attached to a Confluence page, and a page whose space lookup fails.
Chunk and Embed
Documents are split at their headings into chunks of about 500 tokens, usually no more than 800 (a single paragraph longer than that is kept whole), and a long section is split again with a 50-token overlap. Each chunk is embedded using OpenAI text-embedding-3-small (1,536 dimensions) and carries its source document and section heading, so every retrieved answer traces back to its origin.
Retrieve
When a user asks a question, the query is embedded and the closest chunks are retrieved, filtered at the database query level by the user's permission groups. The retrieval never returns a chunk outside those groups, however the question is phrased. Permission is enforced in the SQL, not the prompt.
Respond
Retrieved chunks are passed to Claude (Haiku by default, Sonnet for a question that compares or needs more context) with one hard constraint: answer only from the provided context. If the context does not contain a clear answer, say so. Every response carries numbered citations: each one names the document and links to its source, and in the web app it also shows the opening of the passage. The AI is a librarian. It does not guess.
Learn
Every question left unanswered or answered with low confidence, and every answer someone marks unhelpful, is kept. Once a week the recurring topics are clustered and listed for the knowledge manager as knowledge gaps. The system does not just answer questions. It maps where answers do not exist.
Wednesday Morning at Thornfield
Why Permission Lives in the Database, Not the Prompt
The first version of the retrieval system filtered permissions in the application layer: retrieve documents, then check access. It worked perfectly in testing. The problem was architectural: a carefully crafted prompt injection could instruct the LLM to ignore the application-layer filter and include restricted documents in the response. Moving permission filtering into the SQL query means the database physically never returns documents the user cannot access. No prompt, however crafted, can surface a document that was never included in the query result. The performance cost: 12 milliseconds per query. The security gain: an enforced boundary that cannot be bypassed at the application layer. We rebuilt the retrieval layer on day three of development. That decision has not been revisited.
The System
Three document sources flow through a permission-aware pipeline into a single vector store. Every query is filtered at the database level before the AI ever sees a result. Slack and the web app share one knowledge layer.
The filter is part of the database query, so the model never receives a chunk the person asking cannot see.
Four Decisions That Defined the System
Every knowledge system faces the same tensions: completeness vs. freshness, confidence vs. coverage, automation vs. trust. Here are the four decisions that shaped how this system behaves, and why each was harder than it looks.
Citations are non-negotiable
Every answer carries numbered citations: the document and a link to its source, with the opening of the passage in the web app. We tested a version without citations: cleaner UI, faster responses. Trust dropped measurably in one week. People do not believe answers from a source they cannot verify. The citation requirement also enforces honesty: if the system cannot cite a source, it cannot give a confident answer. The constraint made the AI more reliable, not just more transparent.
Confidence scoring drives documentation behaviour
Every response carries a confidence level, set by how closely the best passage matches the question: high, medium or low. A medium or low answer is marked in colour, so the reader knows to check the source. In month one, 7% of responses were low-confidence. By month three, after the firm used gap reports to fill documentation holes, that dropped to 2.1%. The scoring did not just signal uncertainty. It drove the firm to close its own gaps.
The AI knows when not to answer
The answer prompt allows one source: the passages retrieved for the person asking. When they do not hold the answer, AskBase says "I don't have information about that in the documents I have access to" instead of guessing, and it never speculates. In this scenario that refusal earned more trust from the HR team than any answer did: a question about a colleague, with no document behind it, gets a plain "not in the documents", never a plausible invention. Trust is built by knowing your limits, not by answering everything.
The gap report was the unexpected product
Every unanswered or low-confidence question is kept. Once a week, recurring topics are clustered by embedding-based similarity. At the 90-day review, the gap report showed 14 topic clusters with high query frequency and low documentation coverage. The knowledge manager recognised 11 of them immediately: "That's how we handle conflict of interest disclosures. It's never been written down, we just ask Marcus." The AI's inability to answer became a map of undocumented institutional knowledge. The firm ran a documentation sprint in Q4 using the gap report as the agenda.
Permission in Practice
Same query. Two users. Completely different result sets, because the database never returns documents the user cannot access.
Same query. Same AI. Different permissions, enforced at the SQL query layer, not in the prompt.
What We Learned
Lessons Learned
Permission is architecture, not a feature
We designed the permission model on day one and rebuilt the retrieval layer on day three when we realised the first implementation was not architecturally enforced. Every hour spent getting permission right in the data layer saved ten hours of potential re-architecture, and prevented one category of serious data exposure. In knowledge systems handling commercially sensitive information, permission is not a layer you bolt on. It is the foundation the rest of the system stands on.
The citation requirement changed how the AI behaves
Requiring citations did not just make answers more trustworthy. It made the AI more honest. When retrieved context does not support a confident answer, the system cannot fabricate a citation, so it says it does not know. The citation constraint became an honesty constraint. Users consistently described the system as reliable, which, in this context, meant "it tells me when it doesn't know." That is a harder and more valuable property than answering correctly.
The gap report outperformed the query engine strategically
Knowledge gap detection was built as an afterthought, one day of engineering. In this scenario it is the most strategically valuable output of the system. A leadership team with no interest in "the AI search tool" is intensely interested in a report showing which processes are undocumented, because that report is a documentation plan. The gap report reframes the system from a search tool to an institutional knowledge audit, which is the frame that makes the investment legible above the engineering org.
What I'd Improve
- Deletions on every sync: today a document deleted at the source leaves the index at the next full resync, not the next incremental one
- Knowledge manager dashboard: query volume by topic, citation frequency by document, engagement heatmap by team
AskBase Is the Knowledge Layer
AskBase is the engine that knows. It answers from the firm’s own documents wherever people already ask, in Slack or in the web app, with numbered citations and the same permission filter behind both, so nobody reads an answer drawn from a document outside their groups. Of the four engines, it is the knowledge layer, and the one available now: FlowBase routes incoming work, WatchBase watches what changes outside the firm, and DocBase writes the documents it sends out.
AskBase is the knowledge layer: one index and one permission filter behind every place your team asks.