How do you calculate the Return On Investment (ROI) for Data Quality improvement? We get asked that question almost every day. We know the ROI is significant, but devising a magic formula that fits all scenarios is not. There are too many variables between organizations.
But it is possible to give examples, so I’ll narrow this down to offer a simple scenario and calculations. You can easily apply this thinking into your situation.
I chose ROI for improving CI data, and I’ll look at it through the lens of the Service Desk. My Data Quality Improvement goals are to:
- Reduce incident resolution times.
- Reduce the number of escalations and reassignments.
In my example, improving CI Data quality by a mere ten (10) percentage points produces a staggering ROI of 1436%.
Let me walk you through how I got to this figure.
What is CI Data?
In the context of the Configuration Management Database (CMDB) and Foundation Data, Configuration Item (CI) Data can include things like:
- Identifiers like Name or ID
- Description – what is this CI in human language
- Location — either physical or logical location of the CI
- Criticality or attributes that impact the urgency of any problems
- People and Groups responsible for the CI
- Support/Assignment group, which is necessary for auto-assignments
- Relationships to other CIs and services needed for impact and root cause analysis
For instance, information on a Server’s location, purpose, ownership, management, support, and applications is essential. Surprisingly, this seemingly simple information is often incomplete, outdated, or entirely absent.
One common reason for this gap is that some of this information is not easily discoverable; rather, a human needs to complement what was automatically discovered.

Ignorance is Bliss
We sometimes are told that this is not our concern, and we don’t have this problem because (take your pick)
- We outsourced it and have SLAs in place. This ensures our partner takes care of it.
- Our processes are so good mistakes don’t happen, and there is no evidence to suggest otherwise.
- We do regular cleanup projects, so our data must be good.
You might want to look at this Harvard Business Review article: Only 3% of Companies’ Data Meets Basic Quality Standards (hbr.org). Although a little dated, it is eye-opening.
It is essential to get facts about your data quality. Without visibility into the true state of your data quality, you will not be able to assess your associated risk, cost, and potential impact to your operations.
The Consequences of Invalid or Missing CI Data
So, what happens when CI data is not up to date or is missing entirely? From the Service Desk point of view, these things might happen:
- A server fails. The Service Desk receives a ticket from an automated monitoring system or a user reports a problem.
- The Service Desk agent finds that due to the underlying CI missing data, the ticket does not contain information on who is responsible for the server or what applications run on it.
- The Service Desk Agent then needs to start calling around to figure out whose server this is, what it does, and the potential impact of the failure.
- Only after the agent has enough information can they begin working on the actual problem.
All of this can take hours or even days, and during that entire time, the server is down, and whatever applications rely on it are potentially down, too. The people who need to use that application cannot do their work. Possibly, several people are contacted before the necessary information is found.
Had the data been up to date and readily available, the Service Desk agent could have immediately alerted relevant stakeholders, and work to fix the failure could have begun. Disruptions would have been minimal.
The Consequences Add Up
Worst-case scenarios can be less than pretty. Here’s a real-life example: The 2 million Euro Incident. In that case, a server failure escalated into a costly major incident. The root cause was incomplete CI Data.
Of course, not all incidents are business-critical and can be resolved quickly, even without the CI data. Nevertheless, the numbers add up quickly.
For ROI We Need to Make Assumptions
Cost Per Ticket
I’ll use HDI’s information from 2017, which was readily available. It is from 2017, so the numbers are probably considerably higher today. I’ll round them up to the nearest dollar. According to them, in North America, the cost per ticket varies from $3 to $50, with the average being $16.
So, I’ll use $16 as my Cost per Ticket.
Metric of the Month: Service Desk Cost per Ticket (thinkhdi.com)
Number of Tickets
I will again use information from HDI for this. Their benchmark data is for tickets per user per month, split between industries. This metric ranges from an average of 0.54 monthly tickets per user in the Equipment Manufacturing industry to 1.38 in High Tech.
Without getting scientific about this, I’ll just split the range in the middle and use 0.96 for my example.
Based on that, in my fictional 10,000-employee company, the Service Desk gets 9,600 monthly tickets or 115,200 per year. I’ll use this number in my math.
Cost of Handling Tickets with Incomplete Data
This is a difficult number to come by, but a reasonable approximation can be made from the well-known “Rule of Ten,” which states that “it costs ten times as much to complete a unit of work when the data are flawed in any way as it does when they are perfect.”
In our example, if a ticket has the necessary data when the Service Desk agent gets it, it will cost $16 to handle, but if the data is incomplete or missing, it will cost $160.
For my example, I’ll add a couple of more assumptions:
- 50% of my CI Data is somehow bad. You can quickly get the actual number from a Free Guided Trial with DCM. For free.
- 50% of Service Desk requests rely on CI Data. I’m conservatively guesstimating this based on experience.

Getting Toward the ROI
We already established that the Cost per Ticket is $16 and that the example Service Desk handles 115,200 requests per year.
With this information, I can quickly estimate that about 25% of all requests will require additional work because of missing or invalid CI data. That’s 28,800 tickets. If these tickets did not need extra work, handling them at $16 per ticket would cost $460,800.
But here’s the catch: the cost of handling a ticket that requires additional work, such as calling around to figure out missing location data and then handling everything manually, is tenfold. The cost of resolving these data-missing tickets is an enormous $4,608,000!
Remember that productivity is lost not only by the Service Desk agent and the immediate stakeholders who participate in data gathering. It is also lost by the people who cannot do their work because the ticket takes longer to resolve.
Without going into the discussion about whether 7, 10, or 13 is the correct multiplier, it’s clear that the financial implications are as enormous as they are largely unnecessary.
The Return on Investment
Let’s then get to the Return on Investment. If you invest in Data Content Manager to improve your data quality, the first-year cost for a Starter license is about $40k.
We can reasonably assume that after a year of using DCM to guide your efforts, only 40% of the CI Data would be invalid. That’s 40% instead of 50%, a very achievable outcome in the first year.
Now, the math looks very different:
- Instead of 28,800 requests that require additional attention due to missing location data, you’re down to 23,040.
- The cost of handling those 23,040 tickets is $3,686,400 instead of the $4,608,000 we came up with before the data quality improvement.
That’s $921,600 saved. Almost a cool million!
The ROI is a respectable:
1436%
Return on Investment
And it’s not a one-time return, either. If you get 115,200 service requests this year, you’ll probably get the same next year and the year after, so the improvements you make to your data quality will accumulate over time, and with DCM, you can keep it up to the quality level you choose.
Really?
Of course, examples are simplifications. Nevertheless, this is far from unrealistic. I think it is probably on the conservative side. The data points I used are solid.
You don’t need additional people to run DCM. Instead, those already working with data quality can do their jobs MUCH more efficiently. The cost of training people to use DCM is negligible, and with our new DCM Data Quality Workspace, engaging your data providers to do their part is easier than ever. These are, for example, Business Application owners who don’t often care about data models.
Furthermore, better data quality can allow you to automate simple but manual steps, increasing your ROI even more.
Most likely, many people will be involved in improving and fixing data. With DCM guiding the effort, they will know what to address instead of poking around blindly. You can track and communicate progress with data-based KPIs, not guesswork. When you know you’ve hit a milestone, you can celebrate your wins as you achieve them.
Do it with Your Numbers!
As mentioned in the beginning, we sometimes struggle when asked what the ROI is for Data Content Manager. It is a tricky question to answer on the spot because it depends so much on the organization, cost structure, pain points, current state of the CMDB, and whatnot. Information we are usually not privy to.
I encourage you to do this calculation with your numbers. Better yet, we can help you get some of those numbers in a Free Guided Trial of Data Content Manager. Chances are the outcome will be surprising, justifying a solid business case.
As far as Business Cases go, you might find this article helpful: Build a Winning Business Case for Data Quality Improvement.
Please reach out if you have questions. If you want a baseline for doing your calculations, book a demo with us so we can discuss how to get started for free.












