top of page
Search

Unstructured Data: Tapping Into Inaccessible Knowledge Across the Enterprise

Unstructured data is becoming a growing concern for many organizations. While enterprises invest heavily in collecting, storing, and connecting data to make business-critical information more accessible and valuable, vast quantities of this knowledge are still difficult to use.


The problem? More than 90% of all enterprise data is unstructured and is often buried away in files and applications throughout the organization – for example in presentations, contracts, technical documentation, emails, images, and meeting notes. As a result, less than 1% of this data is currently used in generative AI. Against this background, the traditional focus on data availability is shifting to data usability – with data readiness playing an increasingly important role.


Unstructured Data – The Challenges

To understand the challenges posed by unstructured data, let’s first consider its counterpart, structured data. Data of this kind is organized using predefined formats and schemas. Common examples include customer records from CRM systems, ERP transactions, and operational KPIs. Because of its structure, this information is relatively easy to search, compare, and analyze.


Unstructured data, by contrast, comprises many a wide variety of formats, ranging from documents, presentations, emails, maintenance reports, and technical manuals all the way to images and videos. Not only do these diverse sources lack a standardized data model; their meaning is also often heavily dependent on context.


Fragmentation Erodes Enterprise Knowledge

This means that the challenge facing companies isn’t just high data volumes. It’s also fragmentation across various applications, repositories, formats, and business units. To master this challenge, organizations need to gain an understanding of what information exists; where it resides; who owns it; who can access it; whether it’s current; and how different information is interrelated.


It’s important to recognize that the fact that data exists doesn’t automatically make it usable enterprise knowledge. This characteristic weakness of unstructured data is especially evident in the AI context. Poor data access and poor data quality are two of the leading factors preventing organizations from moving AI initiatives from the experimentation stage into production. For example, an Accenture study shows that only 7% of executives have reached the level of data-readiness needed to scale advanced AI, such as generative, agentic, and physical AI.


Available Doesn’t Mean Usable

As mentioned, data strategy has traditionally focused squarely on data availability: in other words, on collecting, centralizing, storing, and technically connecting information. However, the emerging data-strategy challenge is now data usability, which entails ensuring that information can be identified, understood, trusted, and applied when needed.


Data usability has the following four central dimensions:


  • Context: the meaning of data and its relationship to business processes, assets, customers, and decisions

  • Quality: the accuracy, consistency, and timeliness of data

  • Accessibility: the availability of data to the right users and applications (including the necessary permissions and confidentiality)

  • Traceability: the origin, ownership, and reliability of information


The increasing relevance of these dimensions is reflected in an IBM article, which identifies data quality, lineage, access control, privacy, and compliance as central elements of unstructured data governance.


The core insight here is that simply having data doesn’t necessarily mean that you can create value from it. This is because data isn’t just consumed by employees and analytics applications; it also has to be consumable by AI systems.


Transform and Conquer

Moving from an old data-availability mindset to the new data-usability approach is therefore essentially a shift from volume-driven to value-driven data transformation. This involves transforming fragmented information into accessible enterprise knowledge by applying various criteria, including metadata management and classification, ownership, access management, semantic structures, and data lineage, as well as identification of outdated and duplicated content.


One technological enabler of this transformation is retrieval-augmented generation (RAG for short). This tech has the virtue of enabling LLMs to be connected directly to authoritative enterprise knowledge, without having to retrain the underlying model.


Setting Priorities Based on Value

However, no amount of technology can offset poor-quality data or data that lacks all-important context. So, preparation and governance of the underlying information remain a must. And here, prioritization is key.


It’s not enough to adopt a blanket “make everything AI-ready” approach. Instead, organizations must first identify high-value business capabilities. Then, they work backwards, identifying the knowledge needed for a particular use case and determining what needs to be done to make that knowledge reliably usable.


New Questions for Data Strategy

When it comes to data strategy, companies must take a long, hard look at traditional investment logic. The key question now is whether adding another centralized data platform will create value if the information it contains remains difficult to understand and consume.


Other new strategic questions center on determining which knowledge domains can create the greatest business value. When they know this, organizations can then decide which data should be made usable first. Additionally, there’s the question of how to balance accessibility with security and control. And finally, organizations need to weigh up how much they should spend on platforms and how much to invest in the layers needed to make information usable.


Of course, priorities depend on the specific business context. For example, manufacturing companies will concentrate on engineering and maintenance knowledge while professional services will focus primarily on project knowledge. This development reflects a broader shift in data strategy, away from “How much data can we centralize?” and toward “Which information needs to become usable for which business capabilities?”


Data Readiness: The Bottom Line

Today, organizations don’t necessarily need more data since they already have much of the knowledge they need. The real challenge is now to transform fragmented, unstructured information into knowledge that employees, applications, and machines can use reliably.


Consequently, data readiness increasingly depends on context, quality, accessibility, and traceability, rather than on availability alone. The bottom line is that competitive value shouldn’t be measured by how much information an enterprise holds, but by how much of that information can actually be transformed into usable knowledge and business value.


Any Questions? Any Comments?

If you’d like to discuss how best to tackle your organization’s unstructured data, feel free to contact me. And if you want to share your own ideas on this topic, please leave a comment below.

 
 
 

Comments


  • LinkedIn
  • Mail

Terms of Use
Privacy Policy
FAQ

Copyright © 2024. All rights reserved

bottom of page