In most industrial companies, the knowledge is already there. It is stored in manuals, data sheets, test reports, assembly instructions, and maintenance plans. It lies in folders, on network drives, in PDF archives, and in the minds of a few specialists. When a customer or a colleague has a question, someone has to search for it. The answer exists, it is just difficult to access.
This affects almost every company. According to estimates by IDC and Gartner, 80 to 90 percent of corporate data is unstructured, i.e., in texts and documents instead of maintained database fields. Around 90 percent of this content is never analyzed and remains unused as so-called Dark Data. In the industrial SME sector, this Dark Data is the technical documentation that has grown over decades.
The appeal of AI lies in making exactly this knowledge accessible. However, the path from documents to a reliable assistant involves more than just uploading a few PDFs.
Why uploading PDFs to a knowledge management system is not enough
The obvious first attempt is to feed a few documents into a general AI tool and ask questions. This works for a quick test, but not for productive use, as it often lacks up-to-date information. An uploaded file is a snapshot. If the maintenance plan or the bill of materials changes, the tool knows nothing about it. Provenance is missing. Anyone who provides information in a technical environment must be able to state which source and which version it comes from. Structure is lost. A data sheet relies on tables, units, and limit values, a manual on circuit diagrams and annotated screenshots. Simple ingestion blurs these relationships, and a classic search often only finds what is literally in the text, while diagrams and images remain invisible. And verifiability is missing. A model that formulates freely will deliver a plausible answer even if the underlying fact is missing. How this source integration is achieved technically is explained in the article on hallucinations in AI chatbots.
The path from documents to AI

A burden of documents becomes a reliable assistant when five steps come together.
Capture. First, the sources are gathered: manuals, data sheets, instructions, internal wikis, along with structured data from ERP and master product data. The goal is unified access, not another silo or a new system that requires care and maintenance.
Process. The documents are broken down into meaningful sections and enriched with metadata, such as validity, version, and product reference. This way, the AI will find the correct section later and not just any similar-sounding one.
Link to sources. Every answer is linked to a verifiable section. The language model formulates the response, but the facts come from the document. If no source is found, the assistant admits it instead of guessing.
Keep up to date. Technical documentation changes. The knowledge base must absorb these changes so that the information reflects today's status and not that of the year before last.
Verify and hand over. What the assistant cannot prove is passed on to a human. This controlled handover keeps the information reliable, especially in a technical environment where incorrect information becomes expensive.
What matters most to industry and SMEs
In the industrial environment, decisions with consequences often depend on the answer. Therefore, three points carry more weight than elsewhere.
Technical accuracy is paramount. A confused variant or an incorrect limit value leads to defective parts, downtime, or complaints.
The version status is also decisive. Different documentation applies to a machine from 2015 than to the current model, and the assistant must use the correct one.
And experiential knowledge is at risk of being lost. The DIHK Skilled Workers Report 2025/26 identifies the shortage of skilled workers as the greatest business risk for more than half of the companies. When experienced specialists retire, their knowledge goes with them, unless it is secured in the sources. A well-maintained knowledge base keeps this knowledge in-house. This directly pays off for service inquiries in after-sales, as the article on first-level support in mechanical engineering shows.
From practice: technical support in live operation

A project from technical support shows what this path looks like in practice. A manufacturer of networked industrial devices operates its systems via a central web portal, through which operators in the field access the technology. With each additional device, the volume of inquiries grows. The answers are hidden in long manuals that are rarely opened on-site under time pressure and are sometimes only available in English. Instead of searching, technicians take the direct route via phone or email, and every question ends up in support.
An assistant in the portal addresses three areas. It also understands image-heavy documents: circuit diagrams, tables, and annotated screenshots are unlocked and translated into step-by-step instructions instead of remaining unreadable as images. It merges heterogeneous sources: device manuals, training materials, tickets from the ticketing system, and email histories together form the knowledge base. And it recognizes technical terms in the correct context: domain expressions such as whitelist, activation times, or device ID are assigned to the correct function, even if a user describes them colloquially.
Added to this is the necessary separation of roles: during login, the portal transfers the user's role to the assistant, so that each user group only receives the answers authorized for them. A service technician sees different content than a team leader, and confidential information remains separated between the groups.
The path to get there was short. In the first week, the existing manuals and support histories were unlocked, in the following weeks the knowledge base was connected to the portal without any development effort, and after about four weeks, the first users were using the assistant in the system. The individual solution thus becomes a foundation that can be transferred to other business units.
How Mercury.ai turns documents into an assistant
Mercury.ai brings this path together via the Knowledge Hub. It ingests a company's sources, processes them in a structured way, and links every answer to a verifiable section. In doing so, the Hub also unlocks image-heavy content such as circuit diagrams, tables, and screenshots, processes over 20,000 pages of technical documentation without loss of information, and allows the knowledge base to be updated up to every minute. An orchestra of specialized models recognizes the intent, verifies the source, and only formulates at the very end. This keeps the information bound to the verified knowledge, and the risk of fabricated answers drops significantly. Through the separation of roles, each user group only receives the answers authorized for them, and maintenance is no-code. An overview of the knowledge base is provided on the Knowledge Hub page, while the connection of structured systems is described in the article on CRM and ERP integration.
Frequently Asked Questions
Can you simply upload PDFs to an AI and ask questions?
For a test, yes; for productive use, no. It lacks currency, provenance, structure, and verifiability. A reliable solution links the answers to verified, up-to-date sources.
What is unstructured knowledge?
Knowledge that is hidden in texts and documents instead of database fields, such as in manuals, data sheets, or instructions. According to estimates by IDC and Gartner, 80 to 90 percent of corporate data is unstructured.
How does the AI stay up to date?
The knowledge base absorbs changes to the documents. With Mercury.ai, it can be updated up to every minute, so that the information reflects today's status.
How do you prevent incorrect information from documents?
Through source linking. Every answer is linked to a verifiable section. Without a source, the assistant does not provide information and hands over to a human.
Does this help against the loss of experiential knowledge?
Yes. If the knowledge of experienced specialists is secured and retrievable in the sources, it remains with the company, even when employees leave.
From documents to information
The knowledge in industrial SMEs is there, it is just bound in documents and minds. The path to AI consists of capturing, processing, linking to, and keeping these sources up to date. Then, a burden of documents becomes an assistant that answers reliably around the clock, relieves specialists, and keeps their knowledge in-house.
Would you like to unlock the knowledge from your documents? Talk to us or take a look at the Knowledge Hub.
About the author: Dr. Maximilian Panzner is CTO and co-founder of Mercury.ai. He holds a PhD in computer science from the CITEC Institute at Bielefeld University, where he researched multimodal machine learning and intelligent interaction systems. He has been working on Artificial Intelligence, human-machine interaction, and conversational AI platforms for enterprise use for over 20 years.






