In most industrial enterprises, knowledge is already abundant. It resides in manuals, data sheets, test protocols, assembly instructions, and maintenance plans. It lies in folders, on network drives, in PDF archives, and in the minds of a few specialists. When a customer or a colleague has a question, someone has to search for it. The answer exists; it is just difficult to access.
This affects almost every company. According to estimates by IDC and Gartner, 80 to 90 percent of corporate data is unstructured, meaning it exists in texts and documents instead of maintained database fields. Around 90 percent of this content is never analyzed and remains unused as so-called dark data. In the industrial SME sector, this dark data consists of technical documentation that has accumulated over decades.
The appeal of AI lies in making precisely this knowledge retrievable. However, the path from documents to a reliable assistant requires more than just uploading a few PDFs.
Why uploading PDFs to a knowledge management system is not enough
The obvious first attempt is to feed a few documents into a general AI tool and ask questions. This works for a quick test, but not for productive use, as it often lacks up-to-date information. An uploaded file is a snapshot. If the maintenance plan or the bill of materials changes, the tool knows nothing about it. Provenance is missing. Anyone providing information in a technical environment must be able to state which source and which version it originates from. Structure is lost. A data sheet relies on tables, units, and limit values; a manual on circuit diagrams and annotated screenshots. Simply reading them in blurs these connections, and a classic search often only finds what is literally in the text, while diagrams and images remain invisible. And verifiability is lacking. A model that formulates freely will deliver a plausible answer even if the underlying fact is missing. How this source binding is technically achieved is explained in our article on hallucinations in AI chatbots.
The path from documents to AI

A burden of documents turns into a reliable assistant when five steps come together.
Capture. First, the sources are brought together: manuals, data sheets, instructions, internal wikis, along with structured data from ERP and product master data. The goal is unified access, not another silo or a new system requiring maintenance and support.
Prepare. Documents are broken down into logical sections and enriched with metadata, such as validity, version, and product reference. This ensures the AI finds the correct section later, rather than just a similar-sounding one.
Bind to sources. Every answer is bound to a verifiable section. The language model formulates the response, but the facts originate from the document. If no source is found, the assistant admits it instead of guessing.
Keep up-to-date. Technical documentation changes. The knowledge base must absorb these changes so that the information provided reflects the current status, not that of the year before last.
Verify and hand over. What the assistant cannot verify is passed on to a human. This controlled handover keeps the information reliable, especially in technical environments where an incorrect detail can prove costly.
What matters most for industry and SMEs
In an industrial environment, decisions with serious consequences often depend on the answer. Therefore, three aspects carry more weight here than elsewhere.
Technical accuracy is paramount. A confused variant or an incorrect limit value leads to defective parts, downtime, or complaints.
The version status is decisive. For a machine from 2015, different documentation applies than for the current model, and the assistant must reference the correct one.
And experiential knowledge is at risk of being lost. The DIHK Skills Report 2025/26 cites the shortage of skilled workers as the greatest business risk for more than half of all companies. When experienced specialists retire, their knowledge goes with them unless it is secured in the sources. A well-maintained knowledge base keeps this expertise in-house. For service inquiries in after-sales, this pays off directly, as demonstrated in our article on first-level support in mechanical engineering.
From the field: technical support in live operation

What this path looks like in practice is illustrated by a technical support project. A manufacturer of connected industrial devices operates its systems via a central web portal through which operators in the field access the technology. With each additional device, the volume of inquiries grows. The answers are hidden in lengthy manuals that are rarely opened on-site under time pressure and are sometimes only available in English. Instead of searching, technicians take the direct route via phone or email, and every question ends up in support.
An assistant in the portal tackles this at three key points. It understands image-heavy documents: circuit diagrams, tables, and annotated screenshots are unlocked and translated into step-by-step instructions, rather than remaining unreadable as images. It aggregates heterogeneous sources: device manuals, training materials, tickets from the legacy system, and email histories collectively form the knowledge base. And it recognizes technical terminology in the correct context: domain-specific terms like whitelist, activation times, or device ID are mapped to the appropriate function, even if a user describes them in colloquial terms.
Additionally, the necessary separation of roles is maintained: upon login, the portal transfers the user's role to the assistant so that each user group only receives the answers authorized for them. A service technician sees different content than a team lead, and confidential information remains segregated between groups.
The implementation path was short. In the first week, the existing manuals and support histories were unlocked; in the following weeks, the knowledge base was integrated into the portal without development effort; and after about four weeks, the first users were using the assistant in the system. The single solution thus becomes a foundation that can be transferred to other business units.
How Mercury.ai turns documents into an assistant
Mercury.ai consolidates this journey through the Knowledge Hub. It ingests a company's sources, structures them, and binds every answer to a verifiable section. In doing so, the Hub also unlocks image-heavy content such as circuit diagrams, tables, and screenshots, processes over 20,000 pages of technical documentation without loss of information, and allows the knowledge base to be updated as frequently as every minute. An orchestra of specialized models recognizes the intent, verifies the source, and only formulates the response at the very end. This keeps the information bound to verified knowledge, and the risk of hallucinated answers is significantly reduced. Through role separation, each user group only receives the answers authorized for them, and maintenance is completely no-code. For an overview of the knowledge base, visit the Knowledge Hub page, and the integration of structured systems is described in our article on CRM and ERP integration.
Frequently Asked Questions
Can you simply upload PDFs into an AI and ask questions?
For testing, yes; for productive use, no. It lacks real-time accuracy, provenance, structure, and verifiability. A resilient solution binds answers to verified, up-to-date sources.
What is unstructured knowledge?
Knowledge that is contained in texts and documents instead of database fields, such as in manuals, data sheets, or instructions. According to estimates by IDC and Gartner, 80 to 90 percent of corporate data is unstructured.
How does the AI stay up to date?
The knowledge base absorbs updates made to the documents. With Mercury.ai, it can be updated up to once a minute, ensuring that the information provided reflects the current status.
How do you prevent incorrect information from documents?
Through source binding. Every answer is bound to a verifiable section. If there is no source, the assistant does not provide information and hands over to a human.
Does this help against the loss of experiential knowledge?
Yes. If the knowledge of experienced specialists is secured and retrievable in the sources, it remains with the company even when employees leave.
Turning documents into information
Knowledge in industrial SMEs is already there; it is simply locked in documents and minds. The path to AI consists of capturing, preparing, binding to, and keeping these sources up-to-date. Then, a burden of documents turns into an assistant that answers reliably around the clock, relieves specialists, and keeps their knowledge in-house.
Would you like to make the knowledge in your documents usable? Speak with us or explore the Knowledge Hub.
About the author: Dr. Maximilian Panzner is CTO and co-founder of Mercury.ai. He holds a PhD in computer science from the CITEC Institute at Bielefeld University, where he researched multimodal machine learning and intelligent interaction systems. For over 20 years, he has been working on artificial intelligence, human-machine interaction, and conversational AI platforms for enterprise use.






