Shawn Cordner (00:13.038)
Hey everyone. Welcome to the Amplitude of Tech podcast. I’m Shawn Cordner, chief marketing officer of Amplix. Today I spoke to Michael Smith of Congruity360. I didn’t think he could do it, but he made the topic of data and data readiness sexy. If you don’t believe me, listen to the podcast and see what you think. I enjoyed this one. think you will too.
Shawn Cordner (00:36.27)
Alright, Michael Smith, thank you for joining the podcast today. It’s a tough name to pronounce. I’ve been practicing it all morning. I think I got it right. Is that right?
Michael Smith (00:39.534)
Pleasure’s mine.
Michael Smith (00:47.326)
Yeah, mom and dad pulled out all the stops. My middle name is Joe, so…
Shawn Cordner (00:52.152)
Perfect. All right. So let’s start off with just, you know, not a commercial, but if you want to give us the quick 30,000 foot overview of what Congruity360 does.
Michael Smith (01:02.51)
Yeah, Concruity 360 focuses on unstructured data management. And really what we’re looking at is helping our customers with their number one unstructured data management problems and the issues that occur from having large amounts of data not having in place the ability to manage it properly and what those problems generate from.
From a security perspective, from a data management, from a cost perspective, moving it into, into AI models, which happens to be the number one problem that organizations face when they’re trying to launch AI projects is just having the right data. So we’ve really focused in on a couple of key aspects, putting the right processes in place to match the size and the, problems that you’re having with it. So it’s not a kind of a.
one size fits all, you run everything through the same process, we have multiple steps, and therefore we kind of match up the right process to manage the data more efficiently and do the process more efficiently.
Shawn Cordner (02:11.054)
It’s interesting that we’re in a world now where a company like yours needs to exist. If you go back, you know, 10, 20 years, their businesses had data, right? But unstructured data was not inherently a problem. It was maybe a mess and maybe it would cost a little bit of extra money in terms of infrastructure costs to host all that data. But here we are, we’re living in a data economy. Data probably could be argued as most enterprises most valuable asset. Would you agree?
Michael Smith (02:37.452)
You know, it’s funny you should say that I looked at some historical research analysis and 15 years ago, you’re right. didn’t even, you know, when they ask CEOs what the most important assets are within the organization, data didn’t even get on the list. And now you’d see that right behind, you know, human resources and finances, data is number three on the list or number two. It depends on which industry you’re in.
It really has moved up in importance. think people are starting to understand and take note. And certainly as we’re moving into a more AI centric competitive environment, folks are looking to get advantages and take advantage of their data sources and start using them as intriguing them as assets. So it’s fortunate for us and our industry that people are doing that. And I think you’re also starting to see it.
And some of the frameworks out there, NIST is one of the frameworks that we really follow a lot. And the NIST framework, for example, for AI and also for cybersecurity have moved data management into the forefront. So it’s all good for us, I think.
Shawn Cordner (03:57.13)
isn’t data. think intuitively everybody knows what is data or what data is, but what isn’t data? Because in today’s world and today’s environment, we’re collecting all this data, all these different inputs, all these engagement points and interaction points with the business and the business’s digital footprint. And there’s interactions happening through the contact center and contact centers now are not just people on the phone, but they’re
They’re chat bots and they’re, they’re SMS messages. There’s so many different touch points and there’s so many different, you know, internal processes and workflows and there’s so many human beings, but now there’s also AI agents. so what isn’t data? Like, it feels like everything is data.
Michael Smith (04:40.238)
I think everything is, it is, you know, from a business perspective, I mean, even conversations you and I have in this data, we’re transmitting data back and forth. So when you think about it, human interaction is all about transmitting data to one another. So it could be visual, could be oral, it could be written, you know, all of that is data. But from a data management perspective, there’s really two types of data. There’s two main types of data. There’s unstructured data and there’s structured data.
And your structured data is going to be your databases and your systems that are cranking out data that is organized into standard data-based kind of organizations. Unstructured data is all the other. Now, originally, businesses all focused in on the importance of their structured data. Why? nested heavily in the applications, whether, as you say, it’s ERP systems or whether it’s their financial systems.
taking a look of that was all run on unstructured data. But now, especially, the communication capabilities and how we interact as human beings has been more on the unstructured data management side. So there’s a lot of information that’s captured on the unstructured data side, and that growth has been unbelievable. Unstructured data now accounts for about 85 % of the data that enterprises create on a daily basis.
So when you have 85 % to 90 % of your data is unstructured data and where it’s stored, how it’s stored, who access it, it all becomes really important relative to how a business operates.
Shawn Cordner (06:21.358)
I’m going to take a slight tangent for a second, but are you familiar with the work of Antonio DiMasio?
Michael Smith (06:27.284)
I can’t say that I am familiar with Antonio.
Shawn Cordner (06:30.04)
So he’s a neuroscientist that has done groundbreaking work on emotions. the name of his most popular book is called Descartes Error. And the thesis of the book is that, you know, Descartes said, I think therefore I am. And the is basically I feel therefore I am. And the thesis is that emotions are part of the decision-making process.
even when you think about making rational, analytical, logical thoughts, emotions neurologically pay are a big part of the process of that decision making. And you don’t know it. It’s not emotion in the sense that you’re cognizant of that feeling that you’re having. It’s more an act logical process. But what I love about what you’re you’re saying is you’re putting the human at the center of the data conversation, which I think is so important. So many businesses
especially in a B2B environment, they don’t really think in human terms necessarily. And so what that makes me think about is the data that comes from like sentiment analysis and a contact center, for example. And it’s like, what are we actually measuring there and what can we do with that data? And I think what you’re saying is that there’s gold in a lot of this unstructured data and almost all of the data that we have is actually unstructured. So
I guess that’s a really long roundabout, probably snobby way of getting to the question of what can we do with all this unstructured data? Is it usable?
Michael Smith (08:03.374)
Of course it’s usable, but it’s just like mining. When you think about mining for minerals, how much of that deposit is concentrated in one area and how much is spread out all over the place? speaking of human nature, I always say that human beings have part squirrel DNA in them because they have a tendency to take data and store it wherever they are and around it in multiple copies. So data sprawl becomes…
incredibly important issue to deal with when you start talking about managing data efficiently. So the ability to track data wherever it goes, be able to catalog it, to have the capacity to understand what your data sources are, and then winnow it down. It’s always about culling the data sets down to what’s relevant to the project that you’re working on, what you’re trying to accomplish as a business, what are your business goals.
So whether it’s putting together AI models or just making your business more efficient, is understanding what the goals are and then matching the data sources and the data elements to that and making sure that you’ve got an efficient setup to find that data quickly, easily, and make it still at the same time available to everybody.
Shawn Cordner (09:19.746)
So since we have part squirrel DNA, think one of the things that squirrels do, just to carry your analogy a little bit further, is they take their acorns and they stash them in all these different places, right? And maybe there’s logic to where they put it. Maybe there’s not logic to where they put it. But what we do know is that they don’t always go back and get a hundred percent of the things that they squirrel away. So how do enterprises manage individual people and maybe
broken down processes that don’t put data in a usable or findable or easy.
Michael Smith (09:53.368)
place. Well, I think you’ve nailed the, the, probably the number one issue and problem out there and from data sprawl and data growth and uncontrolled data growth. And that’s the issue around ROT, which stands for redundant, obsolete and trivial data. And the other thing is, you we have, you have turnovers think about, you know, standard enterprise in a business. you have people that work there for a number of years and then they leave the company almost I’d say.
90 % of the time that data becomes unmanaged. It’s just sitting there and it’s sitting out on file shares, it’s sitting in SharePoint, it’s sitting in Fox, it’s sitting all over the place. Why? It’s not anybody’s problem, right? It’s just, you know, that person’s no longer here, out of sight, out of mind, that data sits there. And over the years, you’ve got huge amounts of data sets that are completely unmanaged. And when that happens, you’ve got…
Once again, you could have some really important information that’s in there from a business operational standpoint and also from a governance perspective. And also from a risk perspective. A lot of those files include PII, also includes certainly IP in some instances. So the fact that that’s spread out, it’s unmanaged, it becomes orphan data. Those things really don’t do the business any good.
So the ability to kind of organize that data and make an assessment of that data quickly, because once again, the sheer volumes of data is quite staggering. And the problem, you know, since it hasn’t been addressed in years, you know, and sometimes never in an organization, that culling process can be painstaking. Putting together a good tool set to attack the problem efficiently, handle large amounts of data, take the low hanging fruit, take the easy wins.
the ones that you can identify quickly and move on is always the best practice to start the process.
Shawn Cordner (11:54.306)
So what is the operational impact to the business in having all this data sprawl? think we’re kind of alluding to the opportunity cost, right? There’s an opportunity having this data and certainly not being able to access that data in any kind of systematic way, right? But there’s got to be real costs in terms of storage at the starting point, right? But costs in terms of risk. talk to me about that.
Michael Smith (12:21.602)
Yeah, we typically go in and start doing assessments for companies of their data. We do data assessments. We go in and take a look at their data from multiple different viewpoints. The first one is from infrastructure optimization and understanding that. And what goes into that? Well, you certainly have the underlying storage costs and all the associated costs with storage.
But you also have backup costs, also have network costs, you also have just power and cooling. All the things that go into maintaining a data center or just maintaining a server comes into play. The other piece of the puzzle is that typically everybody leaves it on their tier one storage and they’re backing it up. And you ask yourself when you start doing an analysis of it, and for a while I’ve worked in the area of
enterprise backup. Most of that data is over, you know, two, three, four, seven, ten years. Why are you backing up data that doesn’t change? Number one, the main reason why you’re backing up data is to have version controls to be able to take a look at the data and go recover. And if you have data that’s over two years old, you’re only going to have one version of the data. That’s just the way backups work. So why are you backing that data up?
better off if you move it to a lower cost archive tier that has built in availability, higher availability, you’re going to have better protection of the data. It’s going to be off your tier one storage. You’ve got the ability to call that data back if you need it. So no one can complain that, you know, you’re, you’re, taking my data away. You’re just moving it into a more intelligently managed environment. So that’s, that’s one side of the coin.
From a risk perspective, as I mentioned before, let’s just talk about orphan data. One of the things that happens during cyber attacks is the criminals look for, they look for unmanaged data sources. Why? They can go in and adopt the SIDS and masks themselves to then go ahead and find other sources of data within the organization. Sort of sort of like a mole. Once they get inside of it, then they make it and use it as an identification as they go through the system.
Michael Smith (14:39.372)
So identifying orphan data and things from a security perspective is certainly key. And one of the other things that we also point out is that if you’re leaving these things on your tier one storage, it’s on your network, it’s on your main network. And usually these have Open Read, Write, Max, as permissions and what have you. So it’s available, it’s sitting there, it’s vulnerable. Oftentimes this data includes PII, all kinds of information.
We find data sets where there’s a list of 20,000 security numbers on it in one file. Right? All of that stuff doesn’t belong on your production environment. It doesn’t belong in that environment. It belongs in a protected environment. Pull it off to the side where there’s no access points to it and move it off until you can manage it more efficiently and take it off. So what happens is.
By just going in through and doing basic best practices approach to data hygiene and moving that off of your production environment, you’re able to shrink your attack surface, which is really when you’ll start talking to cyber analysts and start talking about folks who are insurance companies from cyber attacks. Attack surface is a big determination on what your insurance discount rates are going to be. So going in and being able to reduce that attack surface, put it into a
gapped environment where access is completely limited is going to reduce that cyber attack profile for Matic.
Shawn Cordner (16:11.758)
Isn’t that so interesting because I have not heard anyone talk about data being part of the attack surface, right? And that’s, that’s been one of the interesting things in my career to see from a cybersecurity perspective is how the surface area of risk, right? Has grown exponentially. And we saw it so much, you know, during the pandemic where people left the office, right? And so the perimeter of the enterprise network, all of a sudden,
is exploding exponentially because you’ve got all these individuals going home and that just opens up a whole other factor of risk. Now that being exponential because you’re taking one office that had hundreds or thousands of employees and you’re multiplying that risk by that number. Right now we’re talking about data and data is a much bigger set than employees. So that surface area and that risk factor
must be growing by orders and orders of magnitude.
Michael Smith (17:13.524)
Absolutely. You see the research out there on how fast unstructured data is growing. I’ve seen recent estimates where data is tripling over a four-year period. So with that kind of growth, can imagine, once again, you think about all the new applications, sharing applications. As new applications pop up, they all have their applications in their storage area. So each one becomes a silo of information. And that information
They all have different protection levels. They all have different access controls and things in them. So there are gaps that show up as far as being able to access that information. I think being able to identify early on where critical information is and taking steps to secure it. And then once again, if you have multiple copies of the same data, you’re just asking for trouble.
Right. So reduce the number of copies that you have get down to, you know, the minimum and you’re going to really dramatically attack, change your, your tax service and your, your cyber, your cyber vulnerability profile.
Shawn Cordner (18:26.474)
Not only that, but to your point about the social security numbers, you’re also increasing the impact that a breach could have by having all this data sitting around and accessible to bad actors.
Michael Smith (18:40.618)
Absolutely. Once again, most early on in the whole concept of cyber, cyber resiliency and cyber attack defense, it was all focused in on, you know, at the firewall, right? You’re building, you’re building a wall around the organization. You’re putting up your protections against that, but you’re forgetting the number one issue. And the number one thing isn’t about just getting to your defenses. It’s about.
getting access to the data because that’s where the value is. If you look at what a cyber attack represents, they’re going to hold your data hostage, right? And that’s what their goal is. So if you understand that instead of just building up and spending all this time building a great wall, so to speak, around your data and spending all your money and time and energy and effort on building that, I think there are more efficient ways to really
shrink the size of the wall and rarely where you’re going to spend the extra dollars and cents on security on that key data area and moving that off. So, you know, those systems are perfect, but it’s just logical to take what is considered the most important data and put it into the highest security levels that you have.
Shawn Cordner (20:00.878)
Yeah. And building that wall, I mean, you could build the biggest wall out there, but there’s still an open back door, which is the human beings that work for your company and that are part of your supply. And that’s always going to be your biggest risk factor because humans make mistakes and humans fall for things. so bad actors are able to get in through human beings. So while you do need to focus on your defenses, I think you also have to focus on
the mitigation of an incident happening. I would assume maybe you’ll agree with me that the more data that you have, the more complex your recovery is going to be in the case of say, ransomware and you’re trying to restore your data so you can get back into business. The more data you have, the more failure points, the more challenge.
Michael Smith (20:51.502)
Right. And let’s face it, whether it’s backup or whether it’s restoration, you’ve got cycles. And I think by going through and defining what’s critical information, you can triage and put what data needs to be restored first. What’s the most relevant, important, highest value to the organization. thereby determining that, you can set up a restoration process that if you do come under attack and things are going.
what gets restored first is going to be the most mission critical data to you. so out of unstructured data, that is more difficult than, as I said before, when you start talking about ERP systems, it’s pretty defined, you know what it is. And you’re going to do that on the structured data side. But unstructured has been relatively more difficult, but it just requires, once again, identifying and declaring, classifying that data as being important.
tagging it and making sure that you’ve got the systems in place to find it and bring it back as quickly as possible.
Shawn Cordner (21:54.346)
So everything that we’re talking about, we’re dancing around, I think the solution, is data minimization, right? So what does it mean to minimize your data in this context?
Michael Smith (22:03.874)
Yeah, I think you’re going to start seeing, I’ve heard through the grapevine that some of the research analysts that’s going to be one of the key areas that they’re talking about here in 2026 is around data minimization and the benefits of data minimization. And having been in the data side of the field for 10 years plus, it’s refreshing to hear that. Because, when you’ve got the just…
crazy growth of data and most of it being unmanaged. It just leads to really bad issues and problems. So minimizing data, it just requires a different mindset. It requires, you know, once again, doing that assessment, understanding the value of the data and getting rid of the chaff. And that’s just, you know, good data hygiene.
Shawn Cordner (22:51.816)
We had a little pre podcast planning session and you introduced this idea to me and the tension that’s been living in my mind ever since is, you know, use the term rot before, right? So the T and rot is trivial data. And what I’m struggling with a little bit is, like we said, 15 years ago, we didn’t know how important data was going to be.
Five years ago, we didn’t know that AI was going to be as capable as it is today. Use cases are growing and changing every day. We don’t know what tomorrow is going to bring, let alone three years from now. So my question kind of revolves around how do you know it’s trivial versus you just don’t know what the use case is going to be for it in the future. Like you could look at your data and you could say this data here has a clear business use.
This has a legal use. This has a compliance requirement, but the rest of it that isn’t clear today, maybe future opportunity that we’re losing by minimizing the data.
Michael Smith (23:59.532)
Yeah, no, you’re absolutely right. And that there’s a science portion of it and there’s the art of it. Right. And you mentioned before, as we started talking about that, human aspects of things become more important as we go through, especially now that you’re trying to use AI to mimic what human, the human mind is doing that, that becomes important, but there are some certain
things that you can do to start the process. And you’re right, you’re never going to be 100 % accurate or perfect, but there are steps that you can take to help you manage that process. And what do I mean? Just by taking look at metadata, you’re able to make some key decisions about data. And when we start talking about Trivial, we go through with our customers and take a look at, well, what applications are key to your business right now?
and what applications do you think in the future will be and understand that. And there are certain ones that you can agree upon. There are certain applications, number one, that aren’t supported anymore, right? So you can’t even, the data’s sitting there and you can’t even use it because the application, you don’t even have the applications anymore. So just being able to go through and identify applications that are no longer supported or what have you, or not relevant to the business.
You can go through using different pieces of the puzzle there. There’s one data scientist I had talked to before that said he starts every one of his projects taking a look at the DNA of his data and that’s the metadata. So I found that to be a very compelling analogy there because you know, that’s what we preach is taking a look at the metadata from a cataloging and understanding of provenance of the data. You can make certain determinations from that and be able to start the process.
The rest of it requires, you know, contextual analysis. So actually going in and taking a look at the data and starting to find and identify what the key things are. But as you say, that is sitting down and having a conversation with the organization to find out, you know, and agree upon. And that’s not always the easiest thing, but as I said, you can certainly start the process and it takes a…
Michael Smith (26:16.61)
big chunks, big bytes out of data just by doing that cataloging and identification of the data through a metadata analysis. So that’d be my suggestion as a first step.
Shawn Cordner (26:27.274)
I think we touched on this briefly, but if we could spend a little bit more time on it, the process of minimizing your data does have a cost implication in a positive way, right? So can you talk about cost as a forcing function of minimizing your data or as a benefit and outcome of minimizing your data?
Michael Smith (26:44.662)
Yeah, it’s funny. Oftentimes we’ll get pulled into projects and it will be certain stakeholders within the organization. let’s just say we’re going into a specific piece and there might be a merger or acquisition coming in, right? And they’re looking at the new company and they’re trying to determine, you know, what stuff are they going to bring over? What stuff are they going to leave behind?
But going through a discussion with the customer to identify those pieces of the data, the benefits of doing that cleanup and things before you bring it in often pay for the entire project itself, the cost savings. So what we typically do is go in, do a quick analysis, make some suggestions, and based upon those suggestions, model out what the potential cost savings would be on that from a…
you know, from an infrastructure cost perspective, but also say, you know, from a cybersecurity perspective, you shrink the tech surface, kinds of things. So there’s really, you know, a good methodology to put in to taking a look at a project and not only leveraging it for, you know, once again, the actual project itself, what it can do as far as financing, getting cost savings, getting some efficiencies.
scales of efficiencies. And then, you know, it’s usually what you do in those situations is look for the next project. And you typically fly them and you start a project like that. Other folks within the organization say, can’t you do that for my division over here? Can’t you do that for my department over here? Can’t you do that? So consequently, we typically, once we get into an organization and start cleaning up the data, it becomes word of mouth within the organization comes through and people start taking.
Shawn Cordner (28:35.31)
I also wanted to ask about the outcomes with AI because I think intuitively, because the way that most of the public has been introduced to AI is through large language models. And the implication is that more data is better in that kind of scenario, right? You train it, you end up getting a better understanding of how humans talk and interact and what they mean and what the intentions are. But, know, famously right now we’re seeing a pilot to scale gap, right? A lot of people are.
are talking about the challenges of going through a small pilot and actually having that produce an ROI, but also just the intended result, whatever the intended result may have been. I think we’re finding that for these discrete functions that AI is being leveraged for, more data is not necessarily better and that to get better outcomes, you may need to have less data.
Michael Smith (29:27.606)
I would argue that more data is not what you want to do. I would argue that the right data is what you want. You want a representative data set, number one. Because when you think about it, more data is start talking about rot. I start talking about duplicate copies. What does that do to your model?
Right? It’s going to give you bad results because guess what? You’re going to have something that’s overweighted within the model because you’ve got duplicates or you’ve got data that’s outside the range of what you’re looking for. So, I would, I would say, you know, the number one thing is having the right data. And then the other issues when you start talking about large language models and other things such as that, you start talking about.
Okay, what information am I putting in there? Is there PII? Is there sensitive information? Well, you’re putting it in a large language model. You just opened Pandora’s box, right? It’s now been shared. It’s in the model somewhere. You’ve got to be careful. And as I said, the other one problem is we did a survey of six, it through an organization that surveyed 600 enterprises worldwide.
and surveying their CIOs and CISOs. And the number one issue and problem they had with their AI implementations was just getting data, the right data into the model. The cost involved, there’s a reason why Nvidia is the most valuable company in the world is guess what? People are spending a lot of money on processing data. And if you’re wasting it by putting large amounts of bad data in there,
Shame on you. You’ve got to know what it is. So I think a lot of the data scientists that are working on projects, they get it pretty much. it’s folks that are beginning to dabble into it and looking at it from a just dabbling in perspective, they’re going to stub their toes a lot. So I highly recommend that spend a few more bucks on the upside of cleansing that data and getting the right data in there. It’s going to serve you so much better on the outcomes that you get on the back end.
Shawn Cordner (31:46.702)
Is this also a hedge against future operating costs for AI initiatives in the sense that, you you mentioned NVIDIA, these chips, they’re not going to last forever, right? They’re going to be deprecated for years or, you know, there’s better chips that are coming out, right? And so that’s a significant cost that these data centers or enterprises that are doing it on their own are going to incur to get to the next chip.
to realize the next performance power or level because that’s going to be necessary because as you stated before, data growth is exponential, right? It just spinning out of control. So isn’t it a fair statement to say that the less data that you’re processing now and the better your discipline about using data to process in the future, the less your costs are going to increase over time.
Michael Smith (32:41.1)
Absolutely. Cost, time to value. If your models are running shorter, you’re going to see results much faster. So as I said, the costs are minimal on what you’re going to spend on cleaning up versus what you’re going to spend on the backend. So it makes a lot of sense to spend a little bit of time, get some data cleansing going on, get some best practices approach to making sure your data is organized.
advantages when you start getting into the finer points and taking a look at the data and tagging that data with classification context around that data. So as you’re feeding it into the model, you’re also helping the model by giving it some context around that. There’s all these advantages that can happen through just proper prep 101, getting the data.
Shawn Cordner (33:33.526)
Now to drift away from the things that the enterprise has agency over themselves necessarily and kind of talk about maybe market conditions or risk inherent to the market. What are the risks of public model data sharing when it comes to AI?
Michael Smith (33:51.79)
Well, once again, once it’s public shared, it’s available to everybody. So if all of a sudden you put into a large language model, personal information or sensitive information, that’s out there. You have no control over it. It’s gone. It’s as if you went out to the world and said, here’s this data, go for it. So you got to be careful. You got to be careful. And I can tell you, you know, for many years starting, I started up on the wall street.
the, the large players definitely afraid of going into any AI models because they knew they, the amounts of sensitive data that they had, that they just couldn’t allow out, you know, into a large language model or anything. Those, you know, just stop them in their tracks. So there are ways of managing that. And as I said, it’s going through cleaning up the data where you have to, you know,
Take that data out of the pool or do some changes to that data before you implement it. It’s well worth it. I guarantee you, you’ll spend a lot less problems and money, have less problems and certainly from a litigation and liability perspective, dramatically reduce your exposure.
Shawn Cordner (35:09.966)
How can minimizing data affect how the individual employees or people that access that data’s potential for creating problem and leaking data through these public models?
Michael Smith (35:22.266)
There are multiple ways of doing that. We have one customer I’m thinking of as an example, all of their personal data for each one of their employees, they get a final say on how that data is being handled and managed. So we have automated policies that go through and take a look at it and suggest that, but they get a sign off on that. And so they can say yes or no. We then also cleanse the data afterwards.
It’s a matter of making sure that everybody becomes a data steward and in that sense, right? So you’re trying to put together a culture within an organization and putting in place, you know, checks and balances and putting in place that type of structure really goes a long way because if everybody in the organization is taking, you know, taking responsibility for their own data, it just…
works in the ethos for the organization and people attack problems differently.
Shawn Cordner (36:28.11)
Do have any tips for how to impact the culture in that way and make good stewardship of the business’s data, the client’s data part of the culture? Because it’s often hard to get everyone in the boat rowing in the same direction. And I’ve found with other types of projects and implementations, a lot of it is understanding the why and what the are and relating in a way like, you know, in sales and in marketing, right? The first thing they teach you is your language needs to be oriented around
question that the customer or the client is asking themselves which is why do I care about what’s coming out of it? How do you navigate?
Michael Smith (37:06.585)
Well, I think once again, if you can demonstrate from, let’s just say from a financials perspective. So we go in and start doing the following data cleansing and data. We were able to save the company, this amount of money, and all of a sudden we got approval for this part of our budget that affects our lives. Right? So repurposing that, that, that cost center to other projects within the organization, I think is a
great way and a great motivator to build within an organization and ethos of, you’re going to see some benefits of it. And certainly giving people the ability to be their own personal data stewards also engages them and shows them because, you you start saying, well, do you want your personal information out on the, you know, out for cyber criminals and you should get rid of this and, know,
Shawn Cordner (38:01.132)
You fu-
Michael Smith (38:01.452)
Find out that when you start bringing it back and relating it to the individual themselves and how it can affect their life, not only just the organization’s life, it becomes a much easier conversation. think you’ve got more bias.
Shawn Cordner (38:17.45)
So you mentioned NIST before as a framework. I’ve had conversations with cybersecurity experts and a lot of them caution that, you know, leveraging these frameworks is it’s helpful. It’s a tool, but not if you treat it as a check the box kind of approach. Right. So in the context of data management, how would you, how would you encourage technology leaders that are listening to this to
use those frameworks as a tool, where do they need to go above and beyond?
Michael Smith (38:48.222)
Yeah, you’re absolutely right. We’re listening. We’re all into checklists because it makes the job easier. I’ve checked that box. I’m good, out of sight, out of mind. I think by, once again, tying those initiatives into reaching goals, and part of those goals are, let’s say if you have a cost savings, that some of that budget goes back and goes into other projects within the organization that people can benefit of.
I think once again, building that ethos where you’re reinforcing that, you know, this is good for the company, but it’s also good for you as a stakeholder within the organization. I think it’s, you know, building it by department, breaking it down into manageable areas that you can say, hey, if our department is able to show these, show this kind of hygiene, we’re able to represent this kind of savings. It gets shared within the organization, you know, as far as the projects that the department wants to work on.
Shawn Cordner (39:46.882)
There are regulations out there that dictate how businesses can use data and how the people whose data it is that you’re accessing are able to control their own data. And specifically businesses in the U S doing business in Europe have to comply with GDPR. Right. Almost everybody’s going to have to comply with CCPA. There’s, there’s other state level regulations. Right. So how.
Michael Smith (40:15.63)
HIPAA NYDFS all of this,
Shawn Cordner (40:19.31)
So what do you like? How does a business manage all of this? This patchwork of regulations and requirements?
Michael Smith (40:27.022)
It’s funny, I’ve looked at it and I’ve talked to records managers going back through the years and it really hasn’t changed that much. The regulations are all built around specifics about managing things in a thoughtful way. So I think…
Once again, if you can show within the organization that you’ve taken steps to mitigate the risk, first of all, you’re going to be way, way ahead of the game, especially from a litigation perspective. If you can show that you’ve followed a, you know, internal rule to help manage that data, it’s going to reduce any kind of exposure from litigation. You’re still going to have, you know, fines that are associated with, if you’re going to have, let’s say,
release of personal information, let’s say during a cyber attack and you didn’t have any controls in place, you’re going to, you’re going to face funds. So I think it’s a matter of just adhering to a couple of key things. And the key things are number one, classification and putting a classification and identifying information that falls into that class classification. Right? So, you know, this is personal information. It’s got to be managed as personal information and it could be big buckets.
You’re way better off just starting with big buckets, putting in place, making it enforced, then coming up with a taxonomy that’s 25 pages long that falls under this regulation, this regulation, this regulation. You’re better off managing it first at that level and say, okay, what is the key or the longest pole in the tent, so to speak, from identifying.
what is the worst case scenario. So once you do that and you manage that data, you’re now taking that risk profile and made it way more manageable. So I always suggest to organizations, you you got to start somewhere. Don’t try to boil the ocean, break it down into chunks, look and identify what are the key areas that are going to keep you up at night and address those first and then work on the rest afterwards.
Michael Smith (42:48.174)
Because it is complex. mean, when you get into all the rules and regulations, if you want to sit down and look at each rule and regulation in and of itself, it’s like going through an ISO review. It’s pretty detailed. But I say, don’t stop the process because it’s too daunting. Pick your battles, let’s start organizing your data into the bigger buckets. And then there are ways of managing it through the process. Like if you have retention schedules and something’s
You know, you’re not sure whether it’s past a retention schedule. Move it to what we call doing a soft delete. Move it into a protective environment where nobody has access into it. You can take the time and energy, but it’s out of that tax surface and people don’t have access to it.
Shawn Cordner (43:34.926)
Yeah. You take a crawl run or crawl walk run approach, right? Don’t just sit on the sidelines because you can’t do it perfectly. Processed by analysis is not going to get you safe harbor if you do make a mistake, right? But you could at least show that there was intention there and that there were actions there. And this is maybe a turn in the conversation to the next topic, but what you don’t want is for there to be some kind of a breach and for customers that are affected to find out.
that you were sitting on your hands and you didn’t even make an attempt to protect their data because that’s a much different conversation than you did take the right steps or at least the best steps that you could and something still got through.
Michael Smith (44:19.102)
There’s, for example, from a litigation risk, let’s talk about that, getting rid of data. If you can go into a hearing and you got rid of data and they’re saying you got rid of the data incorrectly, if you have a policy and you followed that policy and it was presented and that policy in and of itself is pretty good but you still, might have been some data caught up in that, you are in much better position than if you have.
no policy and that happened. can guarantee you can talk to any litigation lawyer that’s going through it from a privacy and also from a… You’re gonna have a much easier time of getting through that to the other side, much less damage to the organization.
Shawn Cordner (45:04.76)
Sure. And you know, the turn that I wanted to make is to trust and board conversations about data, right? Because being a marketing guy and in particular, my background is branding, right? So I look at a lot of things through the lens of a brand and brand equity is built on trust. Customers, your market, they need to be able to trust that the business is going to deliver on the promises that they make and that they’re going to that
that the product or the service is what they say it is and the outcomes are going to be what they say it’s going to be. And, but they also need to trust that when something goes wrong and they call in, they’re going to be supported. And if there’s a problem or the business makes a mistake, they’re going to make it right as best as they possibly can. And I think that what’s unforgivable from a brand perspective and, and really an existential risk in today’s world is.
If you betray their trust by not respecting their data and protecting their data, because everybody knows that there’s a target on the back, you know, air quotes of the data, right? So that’s, that’s one of the first things that I think a business really needs to look at is how are we going to deliver the best defense is possible to the data that we’ve collected from our customers.
Michael Smith (46:31.426)
Yeah, no, there’s, are a steward of their data. They act like one, right? So I completely agree with you. You’re giving the keys to the kingdom. And you know, many times you look at this information and it gets in the wrong hands, you’re really hurting people. So it makes sense to, right? You can’t be paralyzed and not, you know, continue the key functions of your business, you know, because you’re completely just frozen by the scale of.
You know, the responsibility, but you can do what’s reasonable. And what’s reasonable is once again, I’m saying take care of the low hanging fruit first, start a process. When you need more time to get through something, take actions that can really stop, you know, 95 or 99 % of the issues and problems and then work through it to take care of the rest of it.
doing nothing, sitting on your hands, know, or not doing yourselves any favor.
Shawn Cordner (47:34.776)
So through that lens, how do you go about approaching this conversation to the rest of the C-suite, to the board of directors? How do you justify it? Because this is not sexy stuff, right? So if you’re going to spend money, time, resources on something, I feel like a lot of times the board, the rest of the C-suite, they want it be something that’s going to be impactful and visible. And this is not, it’s impactful, but it may not be visible. It’s not.
revenue generating risk mitigation, right?
Michael Smith (48:07.128)
Shawn, it’s funny you say that because I can walk into a C-suite and talk to a CISO and I said, hey, I just reduced your tax service by 70 % your cyber attack service. by the way, I saved you a bunch of money and by the way, I reduced your cyber insurance. Watch the response. That person, that persona is going to have a much different response than let’s say,
to somebody that’s in, know, finance person will certainly appreciate the cost savings, but it’s not. So you need to be, you know, we put together assessment reports and we’ve kind of tried to think of the different personas that are sitting in the room and go through it and highlight, you know, from that data, what’s that data telling me that might be relevant to them. So it’s, you know, there are different ways of taking a look at what you’re doing, but it, you know, the
The folks that are really getting ahead and the folks that are really competitive were the ones that understanding that, A, my data is important and B, that proper management of that data as an asset is good for everybody within the organization. And it starts at the top. So if you can get somebody up at the CEO level, we made levers out of some of them going in with some of our proposals.
And then walking through it. So start a project, you know, show some momentum. And what’s key to it is once again, doing that quick analysis and getting some low hanging fruit to show some, some progress within the problem early on. And once you do that, the projects typically build momentum and you get core buy-in from the C suite because guess what you’re delivering on what you said quick. Once again, trust.
Shawn Cordner (50:01.494)
Yeah, trust. know, risk mitigation, think, always resonates. We’ve come a long way from say, two, three years ago, when the conversation was the board saying we need AI and the CIO is saying, I don’t know what that means. And now I think the board is still saying we need AI, but the CIO knows what that means. And they probably know where they want to go with it, at least in some respects, right? They know where to start if they haven’t already started, which I think most have. So
The maturity has happened at the, on the enterprise side and maybe not on the board side. And I’ve read some data points and I wish I could remember exactly where it was. I want to say it was something from McKinsey, but basically they were making the case that the board is still basically AI illiterate, right. And in general, and so it’s to be kind of an education process. And so I’m starting to think of the, you need AI conversation to be like AI is the destination and there’s two different ways that you can get.
The first way you can get there is by taking the surface streets that exist today. And that’s the tactical approach. That’s I’m going to pilot this. I’m going to buy this, but we’re going to start that or what you could do. And this takes a little bit more time and it takes a little bit more money and a little bit more resources, but you can start an infrastructure project where you build the super highway to AI. what has to go in that super highway is this type of project data readiness. And it needs to be governance.
and needs to be education. And it’s all these not sexy foundational things that in themselves don’t necessarily have an outcome and don’t seem accretive to the business. But what it does is it enables you to get into the fast lane in the future with other projects.
Michael Smith (51:46.712)
Yeah, I think it’s like anything else that we’ve gone through when something’s new and you don’t have a lot of knowledge around it. What do you typically do? Well, you’re the larger enterprises, you go out to a consulting firm to bring them in and have them take you through the project because they’ve got some experience in doing these projects for other companies of your size and they understand what’s required in that.
For those companies that have gone through that and want to start taking it on itself, they obviously have had some history. They obviously have brought the right people in. They’ve identified a chief data scientist and somebody who works underneath the CIO that has been brought in strictly to handle that. And they’re going to manage that and they’re going to build up the right team around the infrastructure and the right team around the tooling and the data management tools around it to go in.
Or as I said, you do an outside consulting group that’s going to come in and provide you with that expertise and that knowledge. we’re early on as far as the maturity of the marketplace and getting projects done more efficiently and off the route. There’s lot of stubbing the toes right now. So you’re hearing anecdotally and also I think of some of the surveys that I’m seeing out there that these projects aren’t necessarily
being successful early on. so people are having second thoughts about it. The bottom line is I think, you know, the cat’s out of the bag. It’s inevitable. We’re all moving towards a more AI requirement within an organization to be successful competing, you know, in today’s marketplace. So I think it’s inevitable. They’re going to have your ups and downs, but I go back to, you know, some really core basic ideas.
Shawn Cordner (53:23.192)
Yeah.
Michael Smith (53:39.458)
You know, I’ll beat the drum until the cows come home. Bad data in, bad results out. So you’re going to have to clean up the data. You’re going to have to have some tools to do it. You’re going to have some people that have to understand the data and what you’re trying to do. It doesn’t matter if you’re doing large language models or using agentic, generative to, to agentic AI, especially on agentic AI. you’re, you’re essentially teaching the endpoints to be
be autonomous and understand that you’ve got to train them right. So if you don’t train them with the right data upfront, you’re asking for problems. you know, it’s just, think, you know, proves the point that you start off just understanding your data is the first step in any of these projects.
Shawn Cordner (54:25.42)
Yeah. And I think it’s probably worth talking about what a right term, let’s call it a data forward organization looks like. Right. So if you’re leading an IT department, what are the right head counts to be hiring for to start this process and to manage this process? know there’s company like yours that can help with pieces of this, right. But there, there’s certainly some level of internal ownership, right? So what does that look like? Do you think?
Michael Smith (54:55.18)
Yeah, we, know, our company specializes in providing the tools to the folks that do it. So, we work with some large consulting practices and things that come in there and smaller, you know, boutique shops and, and it varies, you know, once again, the complexity of the situation, what the, what the project is, they’re very small, very easily defined, you know, projects that,
have big payoffs that folks can implement. And there’s other ones where they want to boil the ocean and it becomes way more complex. But I think what I’m seeing and the results we’re seeing early on is those projects that are focused in on, again, I keep using the term low hanging fruit, is taking a look at what processes that the organization does that are highly repeatable, highly understandable, that lend themselves to, you
We don’t need to put people on this. can put an AI in the place and do this because it’s completely scalable once you get it right to handle that process. it’s, think, and I said that a data scientists are really, I think the key people that you have to in on it. You’re going to be staffing an organization.
Get yourself a really good, and you want to be doing AI, get yourself a really good chief data scientist. They’re the ones that are going to drive and help you put the right people in place to get the project done because they really are thinking the problem through on how to process that data.
Shawn Cordner (56:29.806)
In the cybersecurity space for years now, the story has been that there’s more demand for people with that skill set than there are people that have that skill set. so those people command a premium when it comes to salary and compensation package. Um, but that’s if you can even get them. Right. So that’s why, you know, MSSPs are so valuable to an organization and they’re, you know, security partners.
Michael Smith (56:56.43)
Sure will data scientist. Yeah.
Shawn Cordner (56:58.54)
Yeah. but so, you know, my question is like, do you think that we’re heading in that? Are we there now? Are we heading in that direction with data scientists?
Michael Smith (57:08.47)
I completely think that’s the case. Yes, absolutely. Absolutely. There are so many folks that are trying to get into the, not even at the large organizations, mid-sized companies that certainly see an opportunity to leverage AI to bring value to it. I would highly recommend that you take a look at consulting.
firms around and get a virtual data scientist to come in and handle the project.
Shawn Cordner (57:45.326)
about seems like investing in the people that you already have may be a good strategy as well. And what I mean by that is, you know, the promise or the thread, depending on your perspective of AI is that it’s going to make certain jobs redundant. And I think that extends in the IT department as well. And so most experts that I talk to, they don’t advocate for replacing people necessarily, but augmenting people with AI. wonder, are there, are there.
Is it a good strategy to look at the people in your department now and invest in their future and your future by giving them education and training in data science?
Michael Smith (58:23.982)
Well, I would argue that whatever your job is within an organization, if you’re a manager, you’re trying to do and understand the value of somebody within the organization. Investing in that talent is always key to a company’s long-term success. You build an ethos where we’re building a monster here that’s going to be
tough to compete against. So yes, I think that’s true. if, you know, having an AI advantage is key, then to your organization and what you’re looking to do in the future. Absolutely. You want to keep, you want to keep a great bullpen out there that you can pull people in from, you know, to fill in. yes, building a bench, building a bullpen. Absolutely. And I would argue that’s the same in, you know, whether it’s data science and AI or
or whether it’s design or whether it’s, you know, it’s not really different, right? You want to keep that bench moving.
Shawn Cordner (59:29.134)
any predictions for the next 12, 18, 36 months when it comes to data and AI that you want to go on record and be held accountable for?
Michael Smith (59:39.918)
I think that, you know, going through boom bus cycles within, you know, within there’s inevitably going to be a slight pullback. I think it’s going to be slight. think, you know, once again, the, the advantages are just too huge that they don’t overcome slight downturns in the marketplace. Nothing goes dramatically up on an old Wall Street or can tell you that. So I think it’s, it’s a little bit cyclical.
But I think, you know, three years from now, you you ain’t seen nothing yet as far as what AI is going to be doing. Yeah, you got smart guys. think Mark Cuban just said that, you know, 40 hour work week is out the door. You know, that we’re going to be, you know, working less and working smarter. But, you
I’m a little hesitant to say, get that bold out there. I just think that you’re going to see continued progress within the marketplace as far as AI adoption. think people are going to get smarter about what they’re using it on and going back instead of trying to boil the ocean. They’re going to start off with those projects at low-hanging fruit, automating within a department, a specific process that’s repeatable.
working on those, and then as you grow and get more sophisticated within the organization, you can branch out. But I think AI is here to stay. I highly recommend people that if you’re not thinking about it, you really should think about it. And as you say, start, start planning for its implementation.
Shawn Cordner (01:01:26.754)
You know, I said earlier that we live in a data economy and we do, but we also live in an attention economy and guys like Cuban traffic and attention, right? And so when they make those bold predictions, they almost never come true. It’s all about, you know, getting the, the tweets and the clicks and all that. But since computers were invented, we’ve been saying the promise of computers and, and the internet was we would get time back and.
you know, we work more, it’s just about an increase in productivity. So I don’t think wherever, I don’t think the 40 hour work week is going away unless it goes the other direction. But, but I do think we’ll get more productivity squeezed out of the people that are putting those 40 hours in.
Michael Smith (01:02:11.722)
yeah, mean, let’s face it. is a, you know, the goal of the company is to return, make returns for the investors and the stakeholders in the company. And obviously if, you know, the alternative is to, you know, cut back on hours. So obviously if you’re going to do that, you’re going to try to lower your costs from a, you know, cost perspective. So I agree with you. I was just making the comment because I consider Mark to be a
pretty shrewd and smart businessman as far as understanding what’s going on within the marketplace. he’s believing that AI is here to stay and it’s going to make a big, huge difference, whether it’s 40-hour workweek or whatever. I’m just pointing to it more as an indicator that the belief in AI, as far as being a component of our go-forward path here, businesses are mature.
Shawn Cordner (01:03:07.64)
I certainly agree with you and Mark about that. Michael, is there someplace that people can find you if they wanted to connect with you?
Michael Smith (01:03:14.956)
Yeah, I’m at Congruity360. My email address is msmith at Congruity360.com. Feel free to drop me a line. Happy to talk to you a little bit about and introduce you to some of the things we’ve been working on.
Shawn Cordner (01:03:29.72)
So I’m Michael Smith. Thank you for your time and expertise. Appreciate it.
Michael Smith (01:03:34.358)
Shawn have a good day.