August 10, 2026

00:21:38

Records Metadata vs Keyword Search: Making Records Findable

Records Metadata vs Keyword Search: Making Records Findable
What Counts?
Records Metadata vs Keyword Search: Making Records Findable

Aug 10 2026 | 00:21:38

/

Show Notes

Records findability starts at the point of creation Trustworthy information has four attributes. It must be accurate, reliable, timely, and findable. Last episode covered the first two. This week, Lee Karas and Maura Dunn take on the other two. The rule does not change. Capture the record where the work happens. Capture it when it happens. Capture it from the people doing the work. As a result, timeliness and findability follow along. Why your file dates cannot be trusted Open a shared drive and sort by date. The dates look authoritative. However, they are not. Sometimes opening a document changes the date. The created date gives way to the modified date. Meanwhile, that modified date may only mean someone glanced at the file. Auto-save makes it worse. Word saves before you rename your copy. Your manual version scheme then breaks quietly. It never runs at all in Microsoft Project or Visio. Even in Excel, it can miss when OneDrive has not finished logging in. Timeliness is proof, not a timestamp An approval date proves timeliness. So capture it deliberately. A workflow engine already knows who drafted, who commented, who approved, and when. Record that context anyway, because property values alone will not carry it. Maura describes a client who did this on paper. Every requirements document carried five or six signature blocks. Each reviewer initialed and dated in ink. Then the team scanned it. An e-signature process gets you the same result today, with less effort. Traceability at every step Business requirement one becomes functional requirements 1.3, 1.4, and 1.5. That trace increases accuracy. The sign-offs at each step add timeliness. Universe search versus precision search Library science calls keyword search a universe search. You give it three words. In return, it gives you everything near those words. For example, ask for "training" and you also get "learning." The result set grows. The usefulness does not. Instead of the whole company, you now face a couple thousand documents. A precision search works differently. First you set the file type. Next you name the business process. Then you name the contract or the facility. Now the found set is small enough to use. Metadata is what makes records findable Four elements matter most: function, process, record category, and trigger date or event. Those are real data, not keywords. Because of that, you can filter on them and build a collection. In other words, you get a virtual folder without a physical file plan. Maura and Lee revisit an old FileNet argument on exactly this point. An IT lead wanted no folders and no file plan. He trusted full-text search to sort out billions of documents. He was partly right, but for the wrong reason. Search, access, and retrieval Findability has three parts. Search identifies potentially responsive material. Retrieval pulls it back so you can use it. Access decides whether you may open it at all. Sometimes users can see that a record exists but cannot read it. Sometimes they cannot see it at all. That choice is yours, so make it on purpose. Episode chapters 00:00 Timeliness and findability: the two attributes we still owe you 00:58 Why the date on the file lies: created versus modified 02:12 Auto-save versus manual versioning, and when OneDrive has not logged in yet 04:35 A tangent worth an episode: password management as a mapping problem 05:34 Capture at the point of creation, then make capture into storage 06:39 What the workflow engine already knows, and what you must record by hand 08:09 Ink initials and signature blocks: traceability from business to functional requirements 09:44 The universe search: three keywords, a million results, ranked by confidence 11:51 The precision search: file type, process, and facility narrow it down 12:58 Process metadata tags every file behind the scenes 14:27 The FileNet argument: "we will just use search" 15:24 Function, process, record category, and trigger as real metadata, not keywords 17:19 Contract records split: construction and non-construction retention 18:47 Collections as virtual folders, and knowing your repository before you commit 19:46 Search, access, and retrieval: three parts of one problem 20:39 Should users know a record exists if they cannot open it? Closing What Counts is produced by TrailBlazer Consulting, LLC and hosted by Lee Karas and Maura Dunn. Learn more or reach out directly at [email protected]. Explore compliance-ready corporate training programs at the TrailBlazer Learning Academy. Read more from Maura at Maura's Substack. Music by Jason Blake. Full disclaimer.
View Full Transcript

Episode Transcript

[00:00:00] Speaker A: Hello, welcome to what Counts. Every organization hides a story in their data. This is a podcast that digs into the governance problems people inherit, ignore and discover too late. I'm Lee and as always, this is my co host, Maura Dunn. Maura, last episode we talked about creating data close to the source as possible. Capture the records where the process makes it. Not three weeks later somebody writes it down in a report, but as the. As it's process processed. So you told a great story about back file characterization, about showing your homework as well. But you did something at the top of the episode that I want to, that I want to carry forward, and that is you said that data has to be reliable, has to be accurate, and it has to be timely and you have to be able to find it as well. So I think we should cover those items. The timeliness and the findability. [00:00:58] Speaker B: Yeah. So I think we did talk about reliability and accuracy and how the closer to the creation point you capture the record, the higher reliability and accuracy is. Timeliness and findability also follow on from that same principle of capture it when you create it, capture it by the people who create it, and capture it at the point of creation or when it comes into your office, your business. So timeliness. So this is when you're looking at a whole. The challenge we have sometimes is when you look at a lot of files just sitting out there on a shared Drive or a SharePoint site or a Google Drive and they've got dates on them and you can sort by date, but honestly the dates are unreliable when you just look at a date in a Properties in a, in an Explorer window or in a Google Drive search or something. Because sometimes when you open a document, the date changes and instead of being the created date, you get the modified date. And that modified date might just mean you opened it, not necessarily that you changed something, but you can't tell. And then you will have the versioning that manual versioning is such a challenge because people rename things, only sometimes they don't. And the autosave feature in Microsoft Word actually leads to a lot of things that get changed and not renamed. Because you're 10 minutes into it, you're like, oh, I forgot to make a copy, I forgot to change the name. Which is meant to be a good collaborative tool, like it's auto saving. You're not going to lose changes if something happens to the document, but what it does is mess up your manual versioning. [00:02:51] Speaker A: I think that's why I find it aggravating. Autosave. [00:02:55] Speaker B: Yeah, it was A real mindset change when that came out. Then you come to count on it. And sometimes it doesn't work. It doesn't work in Microsoft Project, for instance, ever. It doesn't work in Microsoft Visio, ever. They're just not set up for it. But sometimes it also doesn't work in Office or PowerPoint or Word or Excel, depending on how quickly this is what I've discovered. How quickly you log into your computer and open a file. All of the logging into OneDrive may not have taken place and it thinks you're working offline. So you get to the end of a chain. This happens to me frequently. I open up something first thing in the morning. It's my checkbook spreadsheet. I open it up first thing in the morning, make changes, and then it's not auto saved. Before I figured it out, I kept accidentally creating new versions of this document of my checkbook. I've now figured it out and I always check is it autosave on or not? And it has to do with how quickly OneDrive has logged in. So that's a kind of a, whatever, an aside. Just, just, just a reflection of how, how many things are part of everyday operations. When you're trying to create information and use information, you, you're making changes that you may not realize you're making or you're not making changes that you thought you were making. So we talked about the records, process mapping and really understanding at a detailed level what's the business process that is creating information in order to accomplish something. [00:04:43] Speaker A: Since you gave one tangent, I get to give one tangent. We should do a password mapping process because I'm having serious problems with passwords right now. [00:04:55] Speaker B: That's your daily challenge. That's your today's challenge. Anyway, it's the passwords. Okay, well that's a thought. We will try and come up with a good, good episode around that. And I get to have to do our own mapping first. [00:05:09] Speaker A: I get passwords managers and stuff like that and you could just click through and that's great. But what happens when you change a laptop now? You got to start all over and you got to remember that password that you had there. I just, I'm finding it very difficult. [00:05:24] Speaker B: All right, well, sorry, that's not today's time ability because I don't know the answer to that one. So we're going to have to work on that offline and then we can come back and talk about it. I had an answer for my problem of not autosaving because I've solved That one. So in our records process mapping discussion last time we talked about capturing things at the point of creation or incoming receipt and that that increases their reliability and their accuracy. So it can also increase timeliness. As long as the capture becomes storage. Then when you get to search access and retrieval, which is findability, you're using those same rules. The same rules that you go through for the process are built in to capture and store and index the information. The, the process knows that if it's an automated process, the process knows that this document was created at this point in this process. It knows who did it, it knows who changed it, it knows if it was received from X outside or if it was updated with comments. It knows all of that because that's what process automation engines do. If it's manual, you know those things too because you've laid out the process and you're capturing it at the. But you're going to have to, you're going to have to record that information about this was, you know, this was the approved version or this was the version that was sent for approval and then it came back with changes. And so this version 2 is updated with those changes and then this was the final approved version. You know that if you're managing the process manually or your automation engine knows it, your workflow engine knows it, if it's doing it in an automated way, either way, record that information about how it was created, because that's your context. And that's more than just property values that get generated automatically by Microsoft Office or Google Suite. That's all part of your findability. It's also part of proving your timeliness because it captures the date and time when the, when the approval happened. So you know, this is the approved version and it happened on this date and time. Lee approved it. So that adds to your picture of reliable and accurate information by adding that timeliness piece. And again, you can do it manually, but it's harder. It takes discipline to capture it, but it's how people used to do it. And actually a client we had a few years ago would have a, they had a signature block for every major document that was part of their system life cycle. So we submit a requirements document, we submit a business requirements document and there were five or six signatures that had to happen. Who was the author, who was the original sort of owner of this document and approved it, and then three other levels of approval and they each had to initial end date and they did it in ink because that's who they were. But then we scanned it and we captured it. You could, however, do that through an E signature process and get the same thing. Capturing that Maura drafted this business requirements document and the client signed off on, yes, this is an accurate statement of our business requirements. Got it signed off by everybody. Then Maura took that business requirements document and wrote a functional requirements document and went through the same review process. And you could trace it. You had to trace it. In this case, business requirement one broke down into functional requirements 1.3, 1.4 and 1.5, for instance. So you had traceability, which increases accuracy. And you had the timeliness with the sign offs at each step of the process, from draft to review to approval. And in that case it was a manual process. Much easier if you use a workflow engine. So then we get to the findability. And coming from the world of library science, we talked a few weeks ago about how keyword searches and AI searches are sort of universal searches. The universe search in the, in the parlance of library science, which is find all the things that might meet these criteria. And you've seen it. You do a search on, on Google or on another search engine and it'll tell you, and you put in three words, I don't know, red schoolhouse. And it will say, here's all the ones that had red schoolhouse. This batch had schoolhouse, but not red. And it crosses it out or something like that. So same thing inside your business. If you use the Microsoft Office, you know, kind of built in search engine, it's looking across all accessible Microsoft Office locations, repositories, the ones that you have access to, and looking at keywords, and it'll come up with a million things that have those words that you asked for. And it will list them in their order of confidence, which means this one had all three words. And then we go down to the ones that had two, and then we go down to the ones that had one. Somewhere in there it'll add, well, you didn't say this word, but this is a synonym or this is a same concept. And we think that it's equally relevant. You need to look at it and what you end up with looking for [00:11:13] Speaker A: training and it gives you training. [00:11:16] Speaker B: Yes, or the other way around, but you're looking for training and it gives you learning also. So the return set, the found set, the results of your search just gets bigger and bigger, which is a universe search. However, that's not very useful in some cases if you actually want to find a single document that is the one that was approved for this particular situation, that universe search Maybe it narrows things down a little bit, but not a lot. Now you've got, you know, instead of the entire company's worth of electronic data, you still have a couple thousand documents, documents to look at. So the opposite of a universe search in the world of library science is a precision search. And the more parameters that you are able to give your search engine, the more precise your response is. So not just seven keywords instead of three keywords, but fielded searching where you say I'm looking for a document to start with, as opposed to a spreadsheet or a PowerPoint or a, or a data set or recording or something. I'm looking for a document. So you got file type. I'm looking for a document that was created as part of the contract approval process. That narrows you down to the set of data that's been captured as part of your contract approval process. You might even get to I'm looking for a document that was part of the contract approval process for construction on this facility. And now you're down to a very reasonable set. And that's because you've not only did you map your business process, you also captured all the data points along the way of when a document was created, at what point in the process, what was this process for? And the end result of that business process also stores the output. It might store the input in the same place and in the, so in a behind the scenes way, all of these files are tagged with information about the process that it was part of the process that created it, the actors on it, the contract or the facility or the type of contract. All of those pieces of information are part of your metadata about the process and it relays to, it's applied to the data that is part of the process that makes things much more findable. You don't spend time searching across thousands of documents because the way you've built your business and built your information storage has narrowed it down for you. [00:14:12] Speaker A: So I remember discussing, I won't say arguing, but discussing with an IT individual at one of our past clients, long time ago when we were implementing FileNet and this person said, no, we're not going to create categories or folders of any sort. We're not going to create a file plan. We're just going to use search. And he was confident that search was going to work, even though there could be billions and billions of documents that would be, that would show up as a result of one of these searches. But he was, he was confident that you could narrow it down there and be, continue to Narrow it down until you found the document that he wanted. And I just didn't believe it. [00:15:06] Speaker B: You did not. I remember that well. So the, the issue we had there was you wanted to build a physical file structure, electronic but physical hierarchy, based on the retention schedule and the taxonomy. So function and process and record category and trigger, because those things are critically important to grouping records together and disposing of them in accordance with the retention schedule at the right time. His point that he made very badly, just to be clear, and he may not even have gone all the way to where I think we should go, which is if you have metadata that reflects every one of those things, so the function, the process, the record category and the trigger date or event, then you can use a combination of filtered searches, filters and searching around those pieces of metadata that are real data, not just keywords, but real data, to essentially create a collection that is a virtual folder. And that is true. He either didn't explain it well, which as I recall was part of the problem, but I also think he wasn't sure that you needed all that metadata. He thought the keywords in context, the keywords from the document doing full text searching would meet that need. And that's, I think, not true. I think you end up with more false hits, you end up with less certainty when you count on just searching text. Those pieces of metadata that you're going to add in to your repository that capture the information about what happened to this document through its process. When did it get created, when who created it, who commented on it, when did it get updated, who reviewed and approved it, Capturing every bit of that process and building that metadata. And because you can build the record category, the function and process and record category. So that goes right along with that business process. The business process to approve a contract does not recreate, does not end up with records that belong in the HR categories. The business process to create a contract ends up with records in the contract category, the contract record category. And if you have contracts that are construction and non construction, which is a typical split, because construction contract records have a longer life than non construction contract records, then your business process is going to know, is going to split and say, is this a construction contract, a contract for construction or a contract or a contract not for construction? So your business process is going to help with that categorization into the right record category. And so you're building your metadata, which leads you to, yeah, I can do searching and filtering and create a collection based on all the metadata tags or attributes, all of those words are used somewhat interchangeably. Metadata and tags and attributes. It's different technology to do the same thing, which is apply. Apply some controlled vocabulary to your content so that you can search across things and the. You can create a collection which is essentially a virtual folder. It's as if you were going and searching out and pulling all those documents together and putting them in a stack on your desk. [00:19:01] Speaker A: Yes. What I was going to say, just don't put them in Google Drive or something. Something extraneous. [00:19:09] Speaker B: Yeah. So people do use the Google suite. All of the places where you can hold documents are not the same. There is more structure to some of them and it's easier to build structure into some than others. It's easier to move things around and it's easier to take things out. But that is a subject for a different day. All that I will say about that is make sure you understand the repository that you've chosen. [00:19:39] Speaker A: There you go. All right, good. [00:19:43] Speaker B: Okay, so just find a Bill. We talked about Find a findability and I mentioned search, access and retrieval. They are all sort of aspects of findability. So search is looking at the keywords and looking at the metadata and trying to identify potentially responsive data sets, whether it's documents or something else. Retrieval is actually being able to pull them back and view them or use that data in some way. Access governs retrieval because do you have security access to be able to view this data? You might have, depending on the search engine that you're using, you might be able to see that the data exists, but you can't read it. In other cases, you might not even know it exists. It depends on how the metadata is secured versus how the actual data set is secured. And so again, something to think about as you're setting up those protocols is do you want people to know it exists and ask for access, or do you not even want them to know? So those are all parts of findability. [00:20:54] Speaker A: All right, if you have any questions, please send us an email at info trailblazer.us.com or look us up on the web at www.trailblazer trailblazer us.com check out our learning academy at trailblazerlearningacademy.com that's also very important. Thank you for listening. Please tune into our next episode if you like this one. Please be a champion and share it with people in your social media network or subscribe or like. That would be cool too. [00:21:24] Speaker B: Please ask questions. We would love to answer questions. [00:21:27] Speaker A: This is true. As always, we appreciate you, the listeners. Special thanks goes to Jason Blake created our music. [00:21:35] Speaker B: Thanks, everyone. See you next time.

Other Episodes