Back to Blog

Your Project Needs a Knowledge Base

Max SnijdersLinkedIn
August 6, 2026
Share on LinkedIn
Knowledge BaseAIAgentsProject ControlsData
One rule. Four controlled documents.NEAR-CRITICAL FLOAT THRESHOLD10daysContractor review15daysChange report20working daysTIA templatenoneset per projectRecovery methodWhich one governs the job in front of you?An agent that cannot see the other three will answer anyway.

Ask an experienced colleague on a large project a narrow technical question. Say, what threshold the team uses for near-critical float. Watch what comes back.

Almost never a number. What you get is something closer to: careful, there are a few different answers to that, which deliverable is this for? The contractor review fixes it at ten days, but the recovery method deliberately doesn't fix it at all. And two of the documents saying ten are really the same document under different revisions.

That isn't evasion, it's the answer. A bare number would have been useless, possibly wrong, and you would have had no way of telling which.

Now ask an AI agent pointed at the same organisation's document store. You get a number, confidently, drawn from whichever of those four sources happened to rank highest, with no sign that the other three exist.

Ask a colleague. Ask an agent.“What is our near-critical float threshold?”A colleague who knows the jobAnswers with everything around the factCareful, there are four answersWhich deliverable is this for?Two of those are the same typoWe changed this in MarchThat standard has a newer revisionAsk the process-plant team firstAn agent on your document storeAnswers with the fact aloneTen working days.everything that made the answer checkableThe gap is not recall. It is provenance, and knowing when to ask.

What outlives the fact

A major capital project runs for years and cycles hundreds of people through it. By the time anyone asks why the original logic looks the way it does, the planner who set it has moved on. The engineer who chose the modelling approach is on another job. The estimator who built the cost basis has retired.

What survives is a schedule full of assertions: this activity takes forty days, these two are linked, this one can't start until that approval lands. Every one of them was once a decision, made by someone, for a reason, against some source. Almost none of that reaches the person who inherits it.

So what a project team needs isn't really the information, it is the basis of the information:

  • Why is this duration what it is?
  • What is this scheduling assumption resting on?
  • Who decided we'd model it this way, and against what?
  • Which revision of the standard were we holding when we set this?
  • What did this cost estimate get built from?

A fact without its basis is close to worthless, because the only two things you can do with a bare assertion are believe it or ignore it. You cannot check it. And on a ten-year job, checking is the whole game.

Add that up and what a project needs is lineage, which is a different thing from storage and mostly not what the storage was designed to give.

Supposedly a solved problem

Project controls is not short of software: document management systems, project inboxes, scheduling platforms, collaboration environments like Oracle Primavera Cloud. Every one of them exists to hold structured information for a project team, and the good ones are very good at the job they were scoped to do.

They share three problems, and the third one is new.

The estate is piecemeal.The meeting is in one system, the drawing in another, the schedule in a third. The decision is in an email thread and the reasoning behind it is in someone's head. Each tool is competent inside its own boundary and blind immediately outside it.

Nothing cross-references. No link runs from the schedule assumption to the meeting where it was agreed, or from the modelling choice to the standard it came from, or from the estimate to the benchmark behind it. Every tool stores its own artifacts; none of them store the relationships between artifacts, which is where lineage actually lives.

None of it was built to be read by an agent.These systems assume a reader who already has context and is going to look something up: someone who knows this project, knows which folder is the real one, knows that the file marked "final" was superseded in March. That assumption held for thirty years. It broke the moment the primary reader became something with no context at all.

Four things an agent is missing, not one

These get treated as a single problem, and they have different fixes.

  • Common sense. The tacit judgement about what is plausible on a job like this one.
  • Industry expertise. How this discipline works, as distinct from how it gets described.
  • Project-level expertise. How this project works: its conventions, its history, the quirks it has picked up along the way.
  • Specific factual grounding. Real, retrievable facts about the work, each attributable to a source.

Common sense is the one that gets waved at and rarely examined, so it is worth being concrete. A planner looking at a process plant sees that forty percent of the relationships are start-to-start with lag and thinks nothing of it, because overlapping trades in a plant are normal. The same pattern on a building would be a finding worth writing up. No document anywhere in the organisation says that. It is not industry knowledge either, exactly, since the line between the two cases sits in a particular place for a particular firm on a particular kind of job. That is what common sense turns out to be under inspection: not general reasonableness, but a large stock of situation-specific judgement that nobody thought was worth writing down because everyone in the room already had it.

Only the last of the four is a retrieval problem. Which is why an organisation can index its entire SharePoint, wire up a competent search, and still get confidently wrong answers: it solved one of the four and expected the other three to come along for free.

The test that exposes it

Give a system two sources that disagree and watch what it does.

An agent picks one, silently. It has no mechanism for noticing that the disagreement is itself the finding. It cannot tell which source is more authoritative, cannot separate a real difference of practice from somebody's typo, and has no way to stop and ask.

The colleague does the opposite. The disagreement is the first thing they tell you, because they know that's where the risk is.

That is the gap between what you would expect a colleague to know and what you can reasonably expect an AI to know. It does not close by putting more documents in a vector store. It closes by building the thing the colleague has: a structured, cited account of how this project works, with the disagreements left visible.

The disagreement is the deliverable

We recently built one of these for a large engineering firm, assembled from their meetings, their documents, their schedules, and the automation their own people had written. Around four and a half thousand cited facts. The most useful thing that came out of it had nothing to do with volume.

Reading the sources against each other surfaced thirty-one places where the organisation disagreed with itself. The disagreements were entirely internal, between their own sources. Two internal standards giving different earning ladders. Two revisions of one procedure renumbering the same sheets. Four separate documents setting four different near-critical thresholds, each written by a competent person for a sensible reason.

None of that is a criticism. A body of practice built by that many authors over that many years will disagree with itself in exactly these ways. What matters is that today, those disagreements get discovered by someone getting a wrong answer.

So the design decision that mattered most was that the knowledge base does not pick a winner. Where the sources disagree, every version is published, each labelled with the document and revision it came from, alongside the question that would settle it. The governing rule for anything reading it, human or agent, is that it must never silently pick a side.

That reads like an admission of defeat and it isn't. An agent that can see a conflict can ask about it. An agent that can't will guess, and you will never find out that it guessed.

A wiki page for near-critical float, marked Status: contested, publishing four variants each labelled with its source document and revisionMachine-readableAn agent can refuse toact, and say whyFour documents.Four answers. Allfour are current.

One rule, four current documents, four answers. The page states that it is contested and hands the reader the question that would settle it. Illustrative workspace, not a real firm.

With one hard exception

A defect is never published as a choice. When we classified those thirty-one conflicts, only a handful turned out to be real differences of practice. Sixteen were drift or plain typos: a misspelled node name, a method labelled "subtractive" in one heading and "additive" everywhere else in the same document, a standard attributed to the wrong publishing body.

Those get corrected inline, because publishing both sides of a mistake is worse than publishing nothing. It launders an error into an apparent decision, and the reader has no way to tell it was never a real choice. The wrong strings go into an errata index instead, indexed by the wrong string, so that someone searching for what they actually read still lands on the correction.

Separating a real disagreement from someone's mistake is judgement work, and it is where most of the value sits.

What has to travel with a fact

Once you accept that the basis matters more than the fact, the format follows. Every claim carries its source, and a citation has to do three independent jobs.

One fact, and what has to travel with itNear-critical is total float of ten working days or less.cite: a4f2 @ 00:42:17The labelNames the source withoutdepending on any interface.Still readable by someonewho cannot open it.The anchorLocates the exact passageinside that source.Not the document, but thesentence within it.The momentWhen it was actually said,to the second.Recoverable long aftereveryone has moved on.Drop any one of the three and you are back to a bare assertion.A wiki page whose cited sentence carries a numbered marker, with a Sources footer naming the meeting and the timestamp it came fromEvery sourced claimcarries oneNames the meetingand the second,not a filename

The label, on the page. The claim carries a marker, and the source it resolves to names the meeting and the second, not just the file.

The third of those is where the surprise sits. When we measured where the usable knowledge in that engagement had come from, the largest single source was people talking in meetings, roughly half of all the cited facts. The organisation's controlled documents, the material it had spent years and real money storing properly, produced the smallest share of any source.

A meeting recording open at 42 minutes 17 seconds, with the cited sentence highlighted in the transcriptThe citation seeksto exactly hereThe exact sentencethe page quotes

The other end of a citation. The claim on the page links to the second it was said, not to the document it was filed in.

Nearly a quarter more came from a category that didn't exist three years ago: the automation their own staff had written, and the requests they had typed to it. People encode their real working rules into the tools they build long before they write those rules down.

That is an uncomfortable conclusion. The knowledge your project runs on is mostly not in your document management system. It is being said out loud on a Tuesday, and nobody is keeping it.

Meet the sources where they are

If the highest-value material is spoken rather than filed, then capturing it cannot be a documentation exercise that runs afterwards. It has to run where the knowledge is produced: in the meetings, in the working threads, in the schedules themselves.

That is awkward for anyone whose mental model of a knowledge base is a place documents go to rest. But a repository of what was already written down is, more or less by definition, a repository of the part you already had.

It also demands some honesty about the sources. Transcripts degrade in specific, repeatable ways. In that corpus, timestamps were entirely reliable and speaker names were not, so no direct quote in the finished knowledge base is attributed to a named individual. Text pulled from scanned drawings doesn't simply fail, it fails confidently, returning a table shifted by one row that reads perfectly and says the wrong thing. One drawing package rendered five different client names across its own title blocks.

So the knowledge base carries a page about what not to trust, and it earns its place: every entry on it is something that had already produced a wrong answer somewhere in the corpus, and every entry names the defence. An agent that knows which of its sources lie is worth considerably more than one with a bigger index.

Where this leaves you

We are about to ask AI systems to do real work on projects where a confident wrong answer costs months. They will be good at it roughly in proportion to how well we ground them, and grounding is a question of traceability, not volume: facts that carry their basis, conflicts that stay visible, sources labelled honestly by how far they can be trusted.

The standard to aim at isn't a search box that returns the right file. It is the colleague who has been on the job six years. Ask them something they can't answer cleanly and they don't guess. They tell you why it is complicated, and what they would need to know to give you a straight answer.

An agent answering a threshold question by naming four current documents and asking which deliverable and project class the question is forIt asks.It doesn't guess.Each answer attributedto its documentand revisionThis came from ameeting. It is inno document.

The same question, asked of an agent standing on that knowledge base. It names four answers, attributes each to its document, and asks the two questions that would settle it. The second one, about project class, comes from a remark in a meeting that appears in no document at all.

That behaviour comes from what the model is standing on. The model itself contributes surprisingly little, which is good news: the part that matters is within your control, and for most organisations it has simply never been built.