Corpus first
We work on the complete archive, never a sample. What holds for twenty documents breaks on the real corpus: legacy formats, scans without text, duplicate names across versions.
Industrial organisations keep their critical knowledge in catalogues only a few people can read, in standards that change version, and in the head of whoever has been doing it for thirty years. None of it is written where a system can reach it.
The problem is not information retrieval. It is that the relations between documents were never declared: which procedure governs which record, which reference superseded which other, which instruction became obsolete when a process changed. That structure exists only as expert habit.
We study how much of it can be inferred from the corpus itself, where inference stops being safe, and how a system must present its answer so that a technician can verify it instead of having to trust it. In this domain an answer without a traceable source is unusable: the cost of error is a wrong part fitted or a non-conformity in an audit.
You could buy a search engine and point it at the archive. It has been done many times, and the result tends to be abandoned within months: the system answers confidently questions whose answer it cannot support, and the technician stops trusting it by the third time.
That is why this is treated as research rather than installation. What has to be established before building is where inference from the corpus stops being safe, and that boundary appears in no tool's documentation — it is found by measuring against cases the expert already knows how to solve.
We work on the complete archive, never a sample. What holds for twenty documents breaks on the real corpus: legacy formats, scans without text, duplicate names across versions.
The system takes the repetitive nine tenths and hands the veteran technician the tenth that genuinely needs judgement. Selling it as replacement would be false, and would also guarantee that the technician refuses to help build it.
Every answer carries the path that produced it. The deliverable is not a reference: it is a reference plus its application criteria and its sources.
A customer asks for a part by an old code. That code was superseded by another, which was superseded again. An agent drives a real browser across catalogue sources, follows every substitution and returns the whole chain with its sources; a lighter, cheaper model handles the mechanical navigation.
requested ref.→
superseded→
long chain→
current ref.
A certified quality system produces hundreds of documents across many versions. All are stored; nobody knows which contradict each other or which requirement no document covers. The system ingests the full corpus, indexes it and infers the relations nobody had declared.
The gap panel found concrete, actionable things: records with no procedure explaining their use, and jumps in version numbering that weaken traceability of change control. It also produced the opposite of what was expected — confirmation that no instruction was more recent than the procedure it depends on. Knowing what is not broken also has value before an audit.