The culture of data and analytics is starting to take foot among the general population. Books like "Competing on Analytics" and "Supercrunchers" herald a new era of business whereby every little decision is vetted through intense fact-based scrutiny.
Well, it would seem to be that way. The truth is, most processes within most businesses (even the businesses discussed in the aforementioned books) are managed using crude or imprecise measures. The classic example for me is the new IT project. Such projects routinely introduce new business processes, data elements, and business rules. The success or failure of such projects is typically measured based on whether or not the project delivered on time and on budget. The operational ramifications are rarely considered. However, interestingly post-implementation costs [which account for at least 80% of the overall project's costs] are rarely measured. Most people would say that the main reason for this is that it would require periodic follow-up assessments (i.e. "busywork"), and that the original project team has long disbanded. I partially agree with this, but I would argue that there is a bigger issue here.
People (especially those in senior positions) are loathe to be measured. Ask a VP, director, or manager if she likes the idea of measuring her team's performance and she would say "yes!". Ask the same person if she would like to be measured, and you'll get a long winded answer as to why her performance can't be gauged using numeric measures, and should be based on a "360 review" with all manner of testimony and exhibits. Indeed, today's performance reviews are are more like going on trial, rather than a quantifiable measure of performance. It's okay to subject others to KPIs - just not me!
None of this should come as any surprise. But what is interesting is that the new generation of knowledge workers wants to be measured by KPIs, and they want better KPIs to be measured against. How do I know this? I know this because I routinely talk to people who are on the front line of customer service, and I ask them as many questions as I possibly can. Here is what I've learned folks: The generation that grew up with the Internet (usually under the age of 28), and has been asked to work with a computer and a telephone quickly realizes that their job is being measured by a few simple KPIs, and those KPIs can be quickly gamed. This is a generation that looks for inefficiencies on eBay, that has developed methods for finding free music and videos, that can quickly fact check for discrepancies when BS is suspected. These people are playing massive multi-player on-line games, are on every major social network, and are relentlessly logical and efficient when working with rules-based systems (i.e. companies).
It should come as no surprise that the new generation of customer services reps, and other front-line knowledge workers are quick to find the path of least resistance when approaching their job. This is their comfort zone.
The bad news is that the current KPIs that measure these people, are crude and blunt instruments. Taking customer service as an example. There are two basic KPIs which gets used: The first, "Average Handle Time", or the time it takes to get the customer off the phone. The second, "Number of Call-backs", or the degree to which the customer had his issue "resolved". That's basically it. Of course, most companies perform random audits to keep reps on their toes. While this "boogeyman" style of management keeps the train on its rails, it hardly provides any goals to aspire for. Even a callback can made be for any reason under the sun, and it may even be a satisfied customer calling back to spend more money (this contradiction between reality and metrics is often referred to a the "99 foot man paradox"). As for "Average Handle Time"? Well, I once heard a story about a bunch of call centre reps who had a nice scam going where they would simply hang-up on every incoming call (effectively deflecting the caller to other CSRs). Until they were caught, they were being held up as model CSRs for their lightening efficiency.
So how does one actually measure customer service so that someone "gaming" the KPI is forced to provide competent customer service at a reasonable cost? The answer depends a great deal on the company's missions. However, with today's technology there are a lot of options available, especially given that it's now standard to record each and every inbound call. Furthermore, the latest voice recognition software does an impressive job of recognizing the majority of words and phrases. Focusing on this data set alone, we can start pulling out some interesting KPIs. We would first need to convert these unstructured data to structured data sets. This is probably our most difficult task (on a that note, Bill Inmon's "Tapping into Unstructured Data" is one of the better books on this subject). Once we have a handle on our CSR/customer conversation data, we can start mining for certain terms that would indicate satisfaction or dissatisfaction. I'm aware that it's even possible to capture a customer's mood and emotion through vocal tone analysis. In theory, you should even be able to separate out the CSR's tone and words from the customer's tone and words, for even more fine grained analyses.
I am not saying that it would be easy to establish a KPI to objectively measure customer satisfaction. Rather, I am saying that it's not hard to improve upon our current KPIs.
But what's the point? First off, customer service is generally pretty bad these days. This may even have something to do with CSRs gaming the existing KPIs, especially "Average Handle Time". It also has to do with the complexity of services being offered these days, and the natural frustration that goes along with an excess of business rules that we can't possibly comprehended (there is a greater problem here, that I don't have time to get into in this post). My point in all of this is to say that we are headed towards a culture of analytics whether we like it or not, and if our KPIs have flaws in them, then they will be manipulated against us by our customers and employees. The world we live in is complex and nuanced, and one of our best tools at managing this is through simplification through the use of comprehensive indexes. We just need to get better at designing our KPIs, and ensure they provide us with maximum goal congruence.
---
On a completely unrelated note, I have just rolled off a big project, and am looking for new work. My area of expertise is in data management and enterprise architecture, but I'm also very nuts-and-bolts, and enjoy doing everything from: application support; software development; data modeling; requirements gathering; system sourcing and selection; process architecture; change management; external data procurement and aligntment; data warehousing & BI; metadata management; data governance policy; IT strategy; marketing analytics; and pretty much anything else technology or information related you can think of.
I'm based out of Toronto, but am willing to travel and work anywhere in the world where there's interesting work. I am incorporated, and prefer contract work, but would also consider full time work if there's a good fit.
If you know of anything that you think I might be interested in, feel free to contact me at: neil@hepburndata.com
Tuesday, May 20, 2008
Unstructured Data and the Hunt for the Elusive Customer Service KPI / Looking for Work
Posted by
Neil Hepburn
at
10:37 AM
0
comments
Thursday, April 10, 2008
David Letterman can teach us a thing or two about BI
I've been working a lot these days with data visualizations and presentation reports. I must admit that I've learned a thing or two about how people approach data, both from the IT side and the business side. However, after looking at dozens and dozens of data visualizations and executive reports, I have realized that there are effectively two kinds of reports that you can present to an executive, and we should approach and understand them accordingly.
The first kind of report is what I describe as an "entertaining report". These reports rely heavily on data visualizations, and while they can and should convey information. Their primary purpose is to grab your attention (i.e. entertain you), over and above driving decisions. Since human beings are instinctively visual beasts, we have a soft spot for these types of reports. We rarely know what to do with the information we see in these fancy presentations, but we love it all the same as it speaks to us emotionally. The old saying "seeing is believing" is as true now as it has ever been.
The second kind of report is what I would describe as "decision driving". These reports tend to be bland ordered lists, with numbers. However, these reports are not only the most important in driving decisions, but due to their abstract nature (and lack of understanding of the decision-makers predicament) are very difficult to get right the first time. In fact, these reports tend to be an after-thought since we tend to occupy our imagination with the more wonderous data visualizations, and would rather avoid trying to understand the messy world that the decision making manager has to live in. In fact, I'm sure I've even seen some IT folk sneer at the decision makers for not appreciating their glorious art. I've probably sneered myself at one time.
Going one step further, if we ask ourselves how decisions typically get enacted in business, and look at how people take on decisions, we can see that there is s desire to streamline decision making. Managers are expected as part of their role to make decisions on a regular basis. However, because new decisions represent risk, this in turn leads to stress. So, if we can help managers make better decisions without increasing their stress, this is what we should strive for.
I believe that the top 10 (or bottom 10) list is an excellent framework for streamlining decision making. In fact it's so popular, that it is this tool we use to manage our own lives. I maintain my own to-do lists each day. If I need to go grocery shopping, I always have a shopping list in hand. If I need to get my personal spending down, I take a look at my biggest expenses and attack those in order. In other words, the top 10 list provides a framework for grouping decisions together, and therefore making each subsequent decision easier to tackle. Furthermore, since we already know that list items get easier as we go down the list, the entire set of decisions seems less daunting since we can get into a groove and track our progress.
But let's do a thought experiment to give you a better idea of where I'm coming from. Let's say you were the mayor of Toronto, and you had pledged to reduce crime. You might think to first get a grasp of where all the crime is happening. You hire a consultant to explain this to you. The consultant comes back one month later with an impressive heat-map of the City of Toronto showing in excruciating detail where all the crime hotspots are. As the city mayor you recognize all the neighbourhoods, and probably aren't too surprised by what you see. However, you will feel wiser seeing this map as you can now visualize where the crime is taking place (well you might think you can visualize it). Great! Now it's decision time. You need to make some hard spending decisions as to where you want to allocate social spending programs, improve community safety, and boost law enforcement. Is this map sufficient for you to sign into budget these decisions? Perhaps I as mayor could request more heat-maps showing different types of crime like homicide or grand larceny? Maybe an animated time-seriesed map showing the spread of crime might better help? Do you feel confident allocating millions of dollars based on moving blotches on a map, even knowing that those blotches are confiding the truth? Probably not.
I suspect at this point you will want to start generating good old fashioned lists. You might want to see: Top ten neighbourhoods, as ranked by a blended crime index. Or maybe, top ten neighbourhoods, as ranked by velocity of increase in crime over the past 4 years.
There is no shortage of these top 10 lists you could produce, but the whole time you're dealing with unamiguously ranked neighbourhoods, supported by hard numbers. As a compromise, I might say that you could add a simple bar chart visualization to help make some numeric comparisons a bit easier. Either way, you will need to boil things down to a list of some sort, since you will need to verbally articulate the decision you made. What sounds better: "I have allocated an increase spending to: Jane & Finch; Rexdale; and Regent Park, as they currently have the highest indexed crime per capita for the past 5 years standing, according to Statistics Canada". Or, would you rather say: "If you could have seen the map I saw, you would know to allocate funding to Jane & Finch; Rexdale; and Regent Park". Yes, if you're lucky, you might get to hold up the map, but then you would be forced to explain its legend. And if the three neighbourhoods are visually similar to a few other neighbourhoods from a heat-map persepective, then you might be in the awkward situation of squinting your eyes and saying "Well, in my opinion, this blotch looks slightly larger than that blotch". However, if there was a clear winner, then maybe the map wouldn't be so bad? But if there was a clear winner you could state that more clearly in verbal terms.
Show me a data visualization, and I'll show you a top 10 list which does a better job at driving decisions.
However, with all that said, you might be led to believe there are some exceptions to what I am saying. For example, experienced meteorologists are able to make reasonably accurate predictions by visually studying animated satellite imagry of weather patterns. Touche! But I'm not sure if I would even categorize this as BI, since the data never got beyond the video stage into a "fact-based" database. The same goes for military intelligence studying satellite imagry. Once again, the photos are being analyzed as-is for enemy presence.
What I am saying is that the name of the game is to figure out what the most ideal top 10 (actually top 5 might be better due to limits on our capacitative memory) lists are to drive decisions, and you will have saved everyone time, and make managers lives so much easier. However, the hard part is getting into the head of the decision maker. If you cannot understand what the decision maker is confronted with, you will just be throwing darts at a board. Who knows, maybe you'll get lucky.
Thank you David Letterman. You know us so well!
Posted by
Neil Hepburn
at
9:50 AM
0
comments
Saturday, December 01, 2007
The Secret... to Getting Things Done... in Data Management, for Dummies and Idiots... in Three Simple Steps
I've decided to hop on the dumbing down bandwagon for once. But instead of just offering your typically mind-numbing advice, and providing a bunch of pointless examples to prove my point. I'm going to try to address the skeptics head-on.
So what is this "secret", and these "three simple steps" you might ask? I will get right to the point, and then substantiate my arguments for the skeptics. That said, as with anything there are always "exceptions that prove the rule", and if you can point those out (and I encourage you to do so), then I'll try to address those arguments as well. Without further ado, here is "the secret".
- Keep one, and only one copy of your data in an enterprise class RDBMS.
- Normalize your entities to third normal form, or Boyce-Codd normal form.
- Maintain operational data definitions for each data element, ensuring that the definitions. are understood to mean the same thing by all stakeholders.
Problem #1: Keeping only a single copy of the data has the following two major drawbacks:
- Data warehouses need to be kept separate from operational data stores, for performance reasons, storage reasons, structural reasons, and auditability reasons.
- Putting all your "eggs in once basket" is inherently risky, and leads to a single point of failure.
Getting back to reality: First, it is possible with all major RDBMSs (i.e. DB2, Oracle, and SQL Server) to limit a user's resource, and protect another user's resources. In other words, it's possible to provide guarantees of service to the operational systems by denying resources to the reporting systems. If resources becomes a regular problem it is now relatively straightforward [through clustering technology] to swap in more memory and processing power. Second, as far as storage is concerned: Storage Area Network [SAN] and Information Lifecycle Management [ILM], and table partitioning technologies have got to the point where, with proper configuration, tables can grow indefinitely without impacting operational performance. In fact, in Oracle 11g, you can now configure child tables to partition automatically based on parent table partition configuration (before 11g you would have to explictitly manage the child table partitions). While we're on the topic of 11g, it is also now possible to Flashback forever. Namely it's possible to see historical snapshots of the database as far back as you want to go. That said, I don't believe this feature has been optimized for timeseries reporting. So, this may be the last technology holdout to support the case of building a data warehouse. Nevertheless, this is something I'm sure this is a problem the RDBMS vendors can surely solve. Third, as far as structure is concerned, it is not necessary, and in fact can be counter-productive to de-normalize data into star/snowflake schemas. I'll address this in my rebuttal to "problem #2". Fourth, auditability is indeed a big part of the ETL process. But we wouldn't need audit data if we didn't need to move it around to begin with, so it's a moot point. If you want to audit access to information, this can easily be done with current trace tools built into all the major RDBMSs.
On the problem of "putting your eggs in one basket". While I can see that this has an appeal that can be appreciated by the business, it's really a facile argument. Simply put, the more databases you need to manage, and the more varieties of databases you have to manage, results in less depth of knowledge for each individual database, and weaker regimes for each database. If you had all your data in a single database, you could then spend more of your time understanding that database, and implement the following:
- Routine catastrope drills (i.e. server re-builds, and restores)
- Geographically distributed failover servers
- Rolling upgrades (I know this is possible with Oracle)
- Improved brown-out support
Problem #2: Keeping data in 3NF or Boyce-Codd Normal form is difficult to work with, and performs poorly in a data warehouse.
These are also valid points. First, I will agree that people employed in writing BI reports or creating cubes (or whatever reporting tool they're using), prefer to work with data in a star or snowflake schema. Organizing data in this format makes it easy to separate qualitative information from quantitative information, and lowers the ramp-up time for new hires. However, whenever you transform data you immediately jeopardize its quality, and introduce the possibility of error. But more importantly you eliminate information and options in the process. The beauty of normalized data, is that it allows you to pivot from any angle your heart desires. Furthermore, things like key constraints constitute information unto themselves, and by eliminating that structure it's harder to make strong assertions about the data itself. For example, an order # may have a unique constraint in its original table, but when it's mashed up into a customer dimension, it's no longer explicitly defined that that field must be unique (unless someone had the presence of mind to define it that way, but once again human error creeps in). Second, as far as performance is concerned, I agree that generally speaking de-normalized data will respond quicker. However, it is possible to define partions, cluster indexes and logical indexes, so as to achieve the exact same "Big O" Order. Therefore the differences in performance are linear, and can be solved by adding more CPUs and memory, and thus overall scalability is not affected.
Problem #3: It is impossible to get people to agree on one version of "the truth".
While all my other arguments were focused on technology, this problem is clearly in the human realm, and I will submit that this problem with never ever be definitively solved. Ultimately it comes down to politics and human will. One person's definition of "customer" may make them look good, but make another person look bad. Perception is reality, and the very way in which we concoct definitions impacts the way the business itself is run. A good friend of mine [Jonathan Ezer] posited that there is no such thing as "The Truth", and it is inherently subjective. As an example, he showed how the former planet Pluto, is no longer a planet through a voting process. Yes, there is still a hunk of ice/rock floating up there in outer space, but that's not what the average person holds to be important. Fortunately, only scientists were allowed to vote, but if astrologers were invited, they would surely have rained on the scientists parade. Okay, so this sounds a little philosophical and airy fairy. But consider this: According to Peter Aiken, the IRS took 13 years to arrive at a single definition of the word "child", including 2.5 years to move through legislation. That said, what advice I can offer to break through these kinds of logjams, and which I've also heard from other experienced practitioners is to leverage definitions that have already been vetted by external organizations. Going back to the IRS example, we now have a thorough operational definition that provides us with a good toehold. Personally, I try to look for ISO or (since I'm Canadian) StatCan definitions of terms, if the word is not clear. Another great place is Wikipedia or its successor Citizendium. Apart from that, all I can suggest is to be as diplomatic and patient as possible.
Am I wrong? Am I crazy? Or does data management need to be complicated. Well, I guess the software vendors like it that way.
-Neil.
Posted by
Neil Hepburn
at
2:34 PM
2
comments
Tuesday, August 14, 2007
Social Networks for the Enterprise
But before getting into my thoughts on social networks and how I feel they might be applied to the enterprise, I'd like to share a little story with you. A couple weeks ago I sat down for lunch at the local food-court down the road. An old neighbour of mine - Murray - who I hadn't seen in a while, saw me and sat down at my table. Within no time Murray (who is a senior manager for a large Canadian insurance company) started to vent about Facebook. Facebook as you may have heard, is the hottest social networking site out there, and in particular is most popular in Toronto with over 700,000 Torontonians, and growing. Murray brought up the fact that Facebook may have a huge valuation of over $10 billion, but is in fact costing companies significantly more than that in lost productivity. While I'd heard about companies (and even the Ontario government) banning access to Facebook, it really dawned on me as to what a time sucker this thing is. I was tempted to bring up the fact that maybe there were other issues surrounding employee managment, and would argue that this is the overarching problem. However, since that conversation I've read that about half of all corporations now ban access to Facebook.
Moving on, MySpace was the first major social networking site to capture the popular imagination. There were sites before this (Six Degrees comes to mind), but MySpace became a hit for the following reasons:
- It was targeted to, and appealed directly to teenagers: Probably the most socially self-concious group that exists. This has changed somewhat due to concerns over sexual predators.
- It was completely open. Anyone could see anybody else's MySpace page without having to register or login.
- It was a platform unto itself. While building a MySpace page is mainly a "fill in the blanks" exercise, users are invited to add "widgets" and from there "pimp out" their MySpace page. Of course this spawned a widget cottage industry, which in turn makes the MySpace platform more desirable to its users.
- It was targeted to college students: Probably the second most self-concious group that exists.
- It was not so open, which made it more conducive to posting private details. Namely, users could feel more confident about posting personal photographs because the security measures were in place to ensure that only certain people ("friends") could see photos and other personal details.
- It included a news feed which allows you to see all your friends updates. This is perhaps the most powerful [and originally controversial] feature of Facebook, and the one feature that has generated the most stickiness.
- It also is a platform like MySpace. However, it's an arguably more powerful platform since the underlying capabilities of Facebook are more robust, especially the security.
In order to understand why this is, you have to ask yourself the following question: What is the nature of relationships in the Enterprise, and how are they different from relationships in mainstream social networks?
In a nutshell, I would say that the answer is thus: Normal social networks are typically defined by relationships that both parties willingly desire. In the Enterprise, relationships tend to be dictated by the Enterprise, and are thus of a utilitarian nature. While it's nice to work with people we're friends with, this isn't always going to be the case. However, if we can make these utilitarian relationships friendlier this is always a good thing. So, I would propose that any social network for the Enterprise be cognizant of the nature of the relationship, but also facilitate warmer connections. In that regard, divulging a certain amount of personal information is not a bad thing, but should be managed with a greater amount of astuteness, which should take its queue from what is normally discussed by the watercooler, or what would normally be posted on a cubicle wall (e.g. photos of spouse and kids).
So, fleshing out the nature of relationships I will describe the following types of Enterprise relationships that I am aware of, and how I think information should be managed with respect to these relationships. Since this is my fantasy, I will assume that the enterprise has Wikified itself, in the manner that I described a few months ago. The basic types of relationships, the types of information that should be accessible through those relationships, and how those information should be secured are as follows:
First: Operational versus Project relationships.
Operational relationships are ongoing and indefinite. Pretty much everyone has a relationship with HR. Furthermore, everyone has a relationship with the helpdesk. In some cases you will want to maintain personal relationships (HR is a good candidate here), and in other cases you will want to maintain a relationship with a proxy (the helpdesk is a good candidate here). For operational relationships, you don't really need to have very much insight into the documents and data that these entities rely on, and for the most part you would just have their contact details, and a few other things that these persons may make public. For example, HR could post (or link to) information about Insurance companies, company dress code policy, benefits, etc. But you don't need to know which IS/IT systems they are using to manage your benefits as this does not concern you.
Project relationships on the otherhand are temporal, but tend to require greater line-of-site to knowledge. So, as opposed to our relationship with HR, where we don't really need to know HOW they do their job, in the case of project relationships this line-of-sight is usually a good thing. As an example: If I'm on a project working with a team of: software developers; quality assurance professionals; business analysts; systems analysts; and project managers. It would save me time to be able to see what they're up to. Speaking in concrete terms, this means I would like to see what documents they are using (i.e. what Wikis they are frequently accessing), what databases they are connecting to, and generally what they are up to (I am also thinking of a Twitter RSS feed here - btw, Twitter on its own has the potential to be an extremely powerful management tool). I don't need to know everything about their life, just everything that they are doing NOW.
Second: Hierarchical relationships. The Enterprise always has been and always will be hierarchical in nature. Yes, we all aspire to the "flat" egalitarian Enterprise, but frankly speaking this simply goes against human nature. It will never happen as long as hairless apes run the world. However, we can manage it. Namely, it should be simple for our Enterprise social network to apply the correct security and privacy settings based on hierarchy. I should be able to see everything my subordinate is up to, but not so much as what my boss is up to. It's all right if she can see what I'm up to though. It sounds a bit cynical, but this is no different from out Enterprises curently function. As for peers, this gets a bit tricky and should be handled on a case by case basis.
Third: Intra-department versus inter-department versus inter Enterprise relationships. I don't have any hard and fast answers here, but this is definitely something that should be considered. Things of course get tricky when you're talking about relationships that go outside the Enterprise. Typically these would be vendor relationships, and typically from a knowledge management perspective, this is by default a one-way street. Namely, the Enterprise should collect information about the vendor, but be hesitant to share anything with them through a social network. While I can see a time where social networks cross over Enterprises, it's hard to say if this is a priority. To be sure, there is operational information that is routinely shared. For example, a shipping company would keep its customers informed about the status of packages and deliveries. But this hardly has anything to do with insight about any particular person within either Enterprise.
This is just a sketch of how a social network could be implemented in an Enterprise, and if nothing else some of the things that an Enterprise architect should be mindful of. At the very least, it should break down barriers of communication, and although I mentioned earlier that hierarchies are inevitable, they also can get in the way and ironically dehumanize us. As a simple start, if more large organizations had personal pages where people could add a few photos, say a few things about themselves, and post links to frequently referenced documents, it would make the place a lot less intimidating, and much easier for new hires or new transfers.
---
On a completely different note, I was contacted by Michael who writes the Data Governance Blog: http://datagovernanceblog.com/
Michael had some nice things to say about my own blog and I am very flattered and appreciative of that. Although I don't blog that often, one of my main goals has been to connect with likeminded individuals out there who see Enterprise Architecture and Data Management as a professional discipline, and who also understand that the discussion is not about Microsoft or Cognos or IBM or any other silver bullet manufacturer, but is something much more nuanced and sophisticated than any of these tech vendors would portray the problem as being. So, I am more than happy to hear from any others out there who see things the same way I do, or enjoys healthy debate.
For my next blog entry, I've got something a bit more abstract - but with real consequences planned. I am partially basing it on a lecture by my good friend Jonathan Ezer.
Posted by
Neil Hepburn
at
7:47 PM
0
comments
Thursday, May 17, 2007
SOA without IT governance = good luck
Before getting into my post, I wanted to mention an interview I read in this month's Wired with Eric Schmidt, CEO of Google. I want to share with you a small excerpt:
Wired: Google’s revenue and employee head count have tripled in the last two years. How do you keep from becoming too bureaucratic or too chaotic?
Schmidt: It’s a constant problem. We analyze this every day, and our conclusion is that the best model is still small teams running as fast as they can and tolerating a certain lack of cohesion. Attempting to provide too much order dries out the creativity. What’s needed in a properly functioning corporation is a balance between creativity and order.
But we’ve reined in certain things. For example, we don’t tolerate the kind of “Hey, I want to have my own database and have a good time” behavior that was effective for us in the past.
Very interesting... Of all the examples the CEO of Google could come up with in terms of governance, is basically data governance. I think this is an excellent thing to mention when developers get in a hissy about how they're using data. Even the almighty Google adheres to a data governance policy, and the CEO is 100% supportive. Which leads into my blog post, about maintaining SOA services. Something tells me that Google probably does a decent job of governing their web services.
Now onto my point...
The SOA revolution is on in full force. It's the shiniest silver bullet to come around in a long time, and to be sure it has some real benefits that cannot be ignored. Unfortunately, I will be surprised if any companies out there that don't already have a strong IT governance in place will be able to succeed in achieving their desired ROIs. Of course slick new technology doesn't need a business case, as most CIOs are shamed into implementing a SOA program even if there is no specific need - it simply becomes "commons sense".
Before launching into my critique I must point out that I am a huge supporter of the SOA approach. Web services, like those offered by Google, Amazon.com, Yahoo!, eBay, and others(check out: programmableweb.com for a comprehensive directory of web services) are without a doubt a standard that's here to stay. Developing future applications using a SOA model clearly makes a lot of sense.
From a corporate IT perspecitve, the SOA value proposition is two-fold: First it allows for re-usability like never before. In this respect, SOA's direct antecedent is software components (e.g. COM components, or EJBs); Second, SOA makes building distributed systems a whole lot easier. In this respect SOA's direct antecedant is a mishmash of all sorts of technology (e.g. message passing, RPC [which ODBC uses], store-and-forward, etc.).
Now here's the rub. If you're going to switch to building things using a SOA approach, you're probably just going to start building services for new applications. Those applications in turn will be funded by projects, which will be managed by a project managers' whose responsibilities are to the success of the project, and not for the success of IT infrastructure. As the PMI likes to remind us: "Never goldplate". Full disclosure: I am PMI certified. Okay, so what does this mean? This means that while it is possible to build re-usable services. In all likelihood, they will be built for a specific application. Fair enough, when the next project that comes around that needs something slightly differently, we can just extend those services, while at the same time extending the value of those services. Not so fast! The project manager on the second project will likely have to decide: Is it cheaper to extend a live service, or just take the original source code, and extend that instead, creating a brand-new service that is all but identical to the original service. Well, in spite of all the best intensions, most PMs will quickly cost out the price to regression test the current application(s) using the existing interface (not to mention the logistical headache) and will take the path of least resistance by building a nearly identical new service interface. Eventually over time what you get is a balkanized set of services which IT will constantly talk about "re-factoring" or "consolidating", but in reality there's very little discretionary money to complete a major project like that. Instead, what will happen is there will be some kind of required change that will impact all services. At that point IT will have to decide whether or not to consolidate or fix each one-by-one. More often than not, it will be the one-by-one fix that you will see. The costs of fixing each of these services will greatly outweigh the original investments to consolidate services, but it will just be a constant headache that cannot be solved without a major infrastructure overhaul which some IT disaster may eventually justify.
You will of course point out that this type of IT sprawl is really just a lack of IT governance. Of course it is lack of governance. The point is: The discipline required to manage the reusability of web services is no different than the discipline to manage the reusability of data, which in turn requires metadata management, which in turn requires solid data governance, which in turn requires solid IT governance.
To sum up: Implementing a SOA strategy, without any success managing data [and hence metadata], is like boarding a ship with an incompetent navigator. Will you get to your destination? Sure, but it'll take you a lot longer, and cost you a lot more.
Posted by
Neil Hepburn
at
10:08 AM
0
comments
Friday, April 06, 2007
Why basic IT services must be commoditized before we move to the next level
I recently got an e-mail from Gord, a former colleague who just came back from a job interview. Gord was lamenting the fact that in spite of all the talk surrounding data stewardship, metadata, and business intelligence. The reality is that most companies are still interviewing people based on product specific technical expertise.
While on one hand it is becoming more and more apparent that “IT failures” are less issues of technology working, and more issues of poor business alignment, companies when hiring are not specifically requesting these skills, or looking into this track record. So, while an interviewer interviewing for a DBA position could ask the questions:
“How do you ensure that the data modeling policy is being followed? How do you deal with non-compliance?” or;
“Have you ever worked in an environment with a focus on metadata management? Can you describe the challenges, and how you dealt with them?”
Instead, the main questions are these:
“Have you ever completed a major database upgrade project?”
“How do you configure Real Application Clusters in Oracle 10g?”
“Describe a robust back-up regime.”
Reading back the questions, it’s clear that the former set of questions are mushy and don’t have clear-cut right and wrong answers. The latter set of questions are point-blank, and while there may be different ways to answer them, the answers can be easily validated.
But there should be other observations from looking at these questions. The answer to the first set of questions should give you an idea of how business minded the DBA is. While the answers to the second set of questions will give you no such insight; but they will tell you how competent the DBA is at physically managing the database.
For the time being, DBAs that can, say, perform a rolling upgrade with zero downtime is quite the hero indeed. But on the same note, why must that DBA be confined to a single company? Isn’t that ability applicable to EVERY company, regardless of their line-of-business. I mean, if you can ensure that behind the scenes your DB is running flawlessly, why are these trades being haphazardly being reproduced in every IT department. And by the way, I’m not just singling out DBAs here, I would say that at least half of all IT roles have little or no direct linkage to business activity.
Of course, we hear about how these roles are becoming less relevant due to increasing automation, but I don’t buy this. Any system that is automated still requires people to monitor them, as well as people to fix them. Simply put, the DBA, the network technician, the Java/VB/C# developer, and all of those other roles which are not in themselves expressions of the business, are still necessary.
Where I believe the reluctance for change is in outsourcing and in sharing of resources across companies.
Ironically, the companies that have a better understanding and a greater need for things like data governance and business rules management are also the least likely to let go of control over these basic services. The main reason for this is stability, security, and privacy, and an overall greater dependency on IT to automate their business.
On the flipside, smaller businesses, especially those that are growing quickly, are more likely to take a risk with a hosted solution. An IT debacle at a small transport company is far less likely to get pasted all over the following day’s headlines than a screw-up at a major bank. Another difference is that smaller businesses that go with a hosted solution may also see it as more risk averse. By sailing on the same ship as dozens or even hundreds or thousands of other businesses, there is the comfort in that there is “safety in numbers”. As sound as this argument is, large companies are simply not as trusting and often see themselves as more important than any of the hosting vendor’s clients, and therefore more important than the even the sum of its clients. In my own opinion I think there is some validity to this argument, so I’ll leave it at that.
With that said, there is still a gaping hole left to fill. Namely, the vast majority of hosting vendors are what are known as ASPs or Application Service Providers. ASPs typically offer hosted versions of popular workgroup applications, such as accounting applications, reservation systems, and so on. While these applications are all configurable, by comparison to a custom built application, they are extremely rigid and can quickly calcify even a small business’ operations by forcing the business to operate in a predetermined fashion.
At the other end of the extreme are what I would describe as “raw” hosted severs. These are companies which will host an Oracle DB for you, and possibly provide some basic DBA assistance. While these services are definitely a step in the right direction there is little preventing you from shooting yourself in the foot.
So I say we are in need of higher level services that are easily configurable, but are abstracted to a business level.
The solution of course is more generic web services. We are beginning to see these pop-up, but from what I can tell, they are still in their infancy. I must admit, I have not done a recent survey, so I can’t tell you if the services I describe below exist yet, but if they do, I urge you first and foremost to bring them to my attention (free plugs), and secondly to review them yourself. Thus, the services I have in mind are:
1. A hosted business rules management (BRM) solution
2. A hosted business process management (BPM) solution, with business activity montoring (BAM), workflow management, transaction management, and global scheduling
3. A hosted relational database solution, with ETL functionality, and metadata management
4. A hosted forms solution (I know a few of these already exist, and I use them)
5. A hosted reporting solution (although I can’t think of any, I’m sure there must be some of these out there)
6. A hosted user directory and identity management solution (I know that these also exist, but am not sure if they had these types of services in mind)
Now, as I just mentioned above, with the exception of hosted forms, and possibly hosted reporting solution, I haven’t seen any of the other hosted solutions on the market. I would say the main reasons for this are:
1. The transactional through-put of web based applications has yet to match traditional solutions.
2. It is not clear how these services would integrate with each other, and more importantly there are no standards for doing so.
For the first problem, of computational resources I’m not too concerned. If the above mentioned solutions are deployed in a grid-computing fashion, and there are multiple applications running on them, then there are tremendous opportunities for optimization. This is simply a problem that will work itself out over time.
The second problem is much thornier. Standards take time to work out, and can often limit the flexibility of what can be done. A more likely outcome is that a large web services hosting company such as a Google or an Amazon.com may release a single packaged suite than encompasses all of these tools. We can see that both of these companies are already positioning themselves in this way with Google taking a more desktop approach and Amazon.com taking a more back-office approach. However, I don’t see either of them taking this space head-on.
I myself look forward to the day where I can architect, build, deploy, and provide high-level support for a full-blown system for a client in Africa or Asia, Europe, or wherever the business may be, without ever having to worry about anything other than the business details to do my job.
One of the great blessings about a career in IT is that it gives you a window into practically every other business out there. The more we can get away from commoditized details and move towards the technical essence of the business, the more varied and interesting our jobs become. Let’s hope these solutions happen sooner than later.
BEGIN SHAMELESS PLUG
I have recently got involved with a pretty cool internet radio project, that I’m proud to unofficially announce. It’s called TUN3R.com and to put it into a nutshell, it’s a next-generation internet radio portal. What makes it different from other radio/music portals is that all stations are laid out in an expansive grid which you can whiz around with a cross-hair to “tune in” immediately to any station.
The technology itself is very impressive, and what I like most is that it reminds me of the good ole’ days where you could just play around with a station tuner and randomly find things you were never expecting to hear. For now though, the product is not fully baked, and we’re going to be radically improving the searching, including the addition of a new type of search which to my knowledge has not been done yet, so that will really push the site to “11”.
Anyway, please check it out at: http://tun3r.com/
END SHAMELESS PLUG
Posted by
Neil Hepburn
at
9:41 PM
3
comments
Saturday, March 03, 2007
Rebooting Repositories: Are Wikis more viable for Metadata, CMDB, Document Management, and other forms of Enterprise Repositories?
I have been grappling with the general problem of knowledge repositories. It’s been driving me nuts. While metadata repositories have been around for a while, new breeds of repositories are emerging at an increasing rate. In particular CMDB repositories (Configuration Management Databases), Enterprise Architecture repositories, Business Rules repositories, and so on are turning up all the time. The problem of course is that:
a) The information in each of these repositories is related in some way, and those relationships are relevant, and are themselves information (i.e. derived facts)
b) The repositories all have their own data models which cannot be easily integrated.
One approach is to build your own repository, and extend it as needed. Another approach is to take an “anchor” repository and attempt to extend it. So, for example taking a CMDB, and extending it to include entities and attributes for a metadata repository. However, both of these approaches require a great deal of effort to build and maintain, and in the attempt to create a cohesive view of knowledge, we invariably get bogged down in the plumbing of the repository itself. An excellent article which exams this problem, is Repositories Build or Buy by Malcolm Chisholm.
What I feel the problem really boils down to is rigid data models that cannot be dynamically changed, and in turn require integration projects. I am a big fan of the relational model, and to my knowledge, it is the only complete data model that exists. Data that has been properly normalized, constrained, and indexed can answer pretty much any question about itself.
While integrating any two relational data models (i.e. repositories) across heterogeneous systems is always possible, it is usually very difficult. While there are numerous reasons for this, I’ll point out two major issues that will never go away:
- Links in the relational model are represented as foreign key to primary key relationships. Such linkages presuppose a single system [the RDBMS] overseeing both the primary key entity and the foreign key entity. Contrast this to the world wide web, where anything can link to anything, and there is no single system enforcing referential integrity. While this would not be acceptable for a “bet-your-business” operational system, when it comes to knowledge management, I think it’s fair to relax the rules a bit in favour of agility.
- Most software developers treat the RDBMS as a “bit-bucket” to store data, and not as a system in its own rite that is capable of managing data on its own. Furthermore, any entities used by the developers are thought of as “black boxes” to only be interfaced through the developers [typically hidden] interfaces. As such, going in and adding even a single column is fraught with peril. In any enterprise, changing a data model in production is typically the riskiest operation you can perform, and runs the greatest cost due to the amount of analysis and regression testing required.
There is also a major marketing problem with traditional repositories: The average person just sees them as obscure “black boxes” that are of the domain of techy geeks. I have argued in the past that data governance and data stewardship coupled with a well structured metadata repository is necessary if you want to achieve data interoperability, and the purest in me will always believe this, but the pragmatist in me also knows that an 80/20 solution that can be sold and implemented is better than no solution at all.
Thus without further ado, I propose that as enterprise architects, IT service managers, and data managers, we seriously consider a Wiki approach to managing and integrating our knowledge. I.e. a Wikipedia for the enterprise. Now, I’m well aware that like anything that’s popular out there in the internet world, someone is trying to apply it to the corporate world. In other words, I don’t think the idea of an enterprise “Wiki” is anything new. However, I feel that people view Wikis in a very narrow light that does not do justice to its potential, and I’d like to point out some alternative ways in which we could marry Wikis to enterprise repositories, like a metadata repository or CMDB.
Wikis in a corporate sense are often thought of as a combination document management system cum message board. It’s a place where you could put a document that could be about a procedure to backing up a server, followed by a tape retention process. Users could go in and edit the Wiki, and time the procedure itself every changed. They could then record what they changed about the procedure in the edit notes. Anyone who is familiar with a document management system knows that this is nothing new, but for the uninitiated, a Wiki is more approachable and easier to digest. I used to work with developing integrations and add-ons for DOCSOpen (the most popular Document Management System of its time), and while I could argue the merits of a document management system (primarily its third party integrations), I would have no problem recommending a Wiki approach if a client was interested. But I digress…
I believe a Wiki could be extended to hold and maintain corporate documents, Metadata, CMDB data, and all other enterprise repository data, if the following shortcomings could be addressed:
- We need to have more powerful editing tools. The current way of editing a Wiki reminds me of when I used to write essays in university using LaTeX. It was always very precise, and you could get beautiful layouts, and once you knew your way around the mark-up language it was very easy to put together slick looking documents. But I had to create a Makefile just to “compile” my documents, and the idea of asking my peers to edit a LaTeX file was not feasible as it was just too techy for the average person. I was always a big supporter of LaTeX since it worked for me, but I acknowledged that it was basically useless for the average person until user friendly LaTeX editing tools came around.
- We need to have more experience and tools to create Directed Folksonomies. A Folksonomy is basically just a taxonomy that has been created by a user community. For example, you could create a classification system for comic books referring to various genres and subgenres. Of course the problem with a Folksonomy is that it expects the person doing the classifying to know what the various genres and subgenres are to begin with and that they are also using these classifications correctly. A Directed Folksonomy on the other hand simplifies this task for the classifier as it allows them to pick and choose the correct genre and subgenre, and ideally it should provide concise definitions of categories and subcategories. This leads me though to the third shortcoming of Wikis.
- We need more granular security. We need to ensure that select parts of Wikis can be edited or viewed by select users and in only select ways. We would also need to ensure that for Directed Folksonomies that only select users (Data Stewards) could create and edit the Folksnomy definitions, but perhaps allow a greater number of people to tag information using those Folksonomies.
- We need autonomous agents that can modify sections of Wikis on their own. Taking a CMDB example, it would be nice to check a single server page to see what servers are currently up, and for how long. That same page could have multiple sections: some sections being edited by people; and other sections that are only edited by autonomous agents. By allowing both people and autonomous agents to edit the same page, we no longer calcify those data models as the agents would always be aware that the full Wiki is not its dominion, and only a section of it is. Compare this to how RDBMS tables are currently treated, and what the consequences might be if we were to add even a single column to a single table.
- We need better audit tools and processes to ensure Wiki integrity. It would be nice for example to ensure that when a metadata element is pointing to a server entity/Wiki, it’s actually pointing to a server entity/Wiki and not just a dead link. While I don’t believe we need to have such integrity enforced by some overruling system (like an RDBMS), it would be nice to have spiders that could crawl enterprise Wikis as an a posteriori batch process that could point out issues that need to be rectified. I feel that this approach would work better, as it would behave the same way should there be a company merger, and could quickly and effectively assist in integrating both sets of knowledge, without ever actually preventing such an integration from happening due to overly strict a priori constraints.
- We need better reporting tools to allow BI-style reporting. Although I’m suggesting a Wiki approach to entering and maintaining knowledge, there’s no reason why we couldn’t suck this data into a data warehouse for reporting. Such reports could tell us where change is happening the fastest, and by whom, or which areas of knowledge are old and creaky. The possibilities are endless.
Before wrapping up, I’d like to paint a picture of a Wikified enterprise through a simple use case scenario:
You’ve only been with the retail company for a couple of weeks, and have only just got to know the DBAs and a few developers. You’ve been asked as part of a SOX audit to find out where the credit card data is physically stored for its corporate customers, and confirm [or deny] that the server is properly backed-up, and the back-ups are encrypted.
In your past experience you would start making the rounds. You’d be calling people, waiting for responses, following-up with more questions, requesting documentation, and not always getting it. In the end you’ll get your answer, but it could take you the better part of a week just to track down the right people and get them to locate the correct documentation.
In my dream Wikified enterprise, things would instead go like this:
- You go to the Wiki portal where you search for “credit card”.
- You find various results for pages with credit cards, but near the top you find a “credit card” data type page. You decide to click on it.
- The page brings up a definition of the credit card type, and includes links to all the various entities that utilize this type. The information on these pages is structured within sections, and has been edited and maintained by knowledgeable data stewards.
- You click through all the entities that have the credit card attribute (there would be links from this page). You scan the definition of each entity until you find those entities that pertain to corporate customers. The definitions of the entities should provide this information, or at least provide links to other Wikis which would provide this information.
- You have now located all entities that hold corporate credit cards. You now need to determine where these entities are physically stored in production. The entities themselves would have links back to which DBMS they are stored in.
- You then click on the DBMS Wiki link for each entity to find out where they’re stored. From here you have located three DBMSs: a data warehouse; an ODS; and a third DBMS.
- For each DBMS Wiki, you click on its link to find out more about it. One of the sections has links to the physical hardware that the DBMS runs on. You then go to this server Wiki. You read up on the server’s security and make notes. The server Wiki has a link to the data centre Wiki. You then click on the Data Centre Wiki to read about it, where it’s located, and its security policies. You bookmark this Wiki, while taking notes.
- You go back to the DBMS Wiki to read about the back-up and retention policy, which is contained in its own section. You note that there is no mention of whether the back-up tapes are encrypted or not (even after checking the respective back-up server Wikis), and decide to call the person who last edited the back-up section. After getting in touch with this DBA she informs you that the tapes are in fact not encrypted. Good to know. No problem, you quickly go into the Wiki and make a note that the tapes are not encrypted, citing where you got the information from.
- You now take the information that you have collected and produce a report. You look back and realize that it only took you a couple of hours to collect and digest all the information, including the call to the DBA. The report took another hour to draft and format, and you feel confident if challenged that you can corroborate your facts. A job well done, in complex environment.
This to me is what the agile enterprise is all about. Clearly we’re still a ways off. However, I’m optimistic that we’ll be there sooner than we think.
There is one last thing I’d like to leave you with on this topic. I have been blabbing on about metadata for quite some time now. While most people “in the know” are quick to agree that metadata is essential to getting to the root of IT failures, it has yet to capture the popular imagination and I fear it never will. On the other hand, Wikis have captured the popular imagination. Both my parents know what they are, and can envision their use in many different ways. I mention metadata and I get blank stares. I mention Wikipedia, and I’m always in for a lively discussion. So, as a pragmatist I feel that even though there are technology hurdles to clear in bringing Wikis to being robust enterprise repositories, I feel that it is a cinch in comparison to convincing people about the merits of metadata. So, the next time you start to talk about metadata or knowledge management, try instead starting off talking about “Wikipedia for the corporation”, and go from there. I bet you will have a much better chance of engaging the person you’re speaking with, whether they agree with you or not. Dialogue is only the beginning.
Posted by
Neil Hepburn
at
11:14 PM
2
comments
Tuesday, January 09, 2007
Implementing Data Governance
If you’ve every purchased a home, (or prepared a Last Will and Testament, negotiated a severance package, or gone through a divorce), it’s likely that you have hired a lawyer to guide you through the process. If you managed to find a good lawyer, hopefully she did a decent job ensuring your best interests were being served. Hopefully she presented you with the key decisions that need to be made, adviced on how to evaluate those decisions, and did not overwhelm you with the minutiae of the law.
You probably also hired a lawyer that was an expert in her field, whether it be real estate law, wills and trusts law, employment law, or family law. This person probably knew her subject pretty well, although she wouldn’t know your specific wishes until meeting with you. She would still have a good understanding of how your situation fits in to the context of the law. To complete the required work, your lawyer may even rely upon other lawyers or legal workers to flesh out all the details and take care of some of the routine work. Continuing with our real estate example, your lawyer would then liaise with the Land Registry Office (that’s what they call the place where deeds are recorded in Ontario) to ensure that the ownership has been officially recognized and that there is a clear and unambiguous record stating your entitlement to the land. At the end of the process, assuming things went as planned, you would be responsible for and take ownership of your new home, while relinquishing ownership of your old home.
In the world of Data Governance things should work in a similar way. However, instead of the client taking ownership of a property, the client takes ownership of a particular Subject Area. Instead of working with a lawyer, the client would work with a Data Steward. Instead of storing your official records with the Land Registry Office, you would instead store pertinent information in a Metadata Registry (often referred to as a Metadata Repository). Finally, there would be a straightforward and unambiguous process to follow, as is the case when dealing with a real estate transaction. Now, I’ve probably gotten a little ahead of myself here.
Questions…
What exactly does it mean to take ownership of a Subject Area? What is a Subject Area for that matter? What is a Data Steward? And what is a Metadata Registry?
A subject area is an area or interest or subset of data concerning something of importance to the enterprise. This could be Customers or Employees or Products or Stores or Financial transactions. To take ownership of a particular Subject Area not only entails taking ownership of the quality of data, but just as importantly entails taking ownership of how that subject area’s Data Elements are defined (a Data Element is just a constrained piece of data, such as a credit card number, or customer address. This is not to be confused with unconstrained or unstructured data, such as a paragraph in a blog). Establishing ownership of a particular Subject Area is probably the most difficult and politically charged aspect of data management. It is also the reason why successful data management professionals must have strong interpersonal skills to succeed. Data management is not as much about technology as it is about PR, marketing, persuasion, and negotiation.
Wikipedia defines a Data Steward as a role assigned to a person that is responsible for maintaining a Data Element in a Metadata Registry. While this is not incorrect, I feel that the role also encompasses the following:
Is assigned to, and has an in depth understanding of one or more Subject Areas (e.g. Customers, Products, etc.).
Has good negotiation skills: This is necessary when attempting to extract information from both internal and external colleagues. The accuracy of information is as good as its source, and Data Stewards must both work hard and use finesse to get people on their side to provide the most accurate and detailed data definitions.
Liaises with other Data Stewards.
Understands the Enterprise’s method of data classification. E.g. the various value domains, value domain types, value domain type classes, etc. (yeah it’s a bit like zoology in more ways than one)
Understand how to relate Data Elements to business and IT entities. E.g. how the customer credit card data element relates to business rules, business cycles, or even which physical servers it is stored on.
Is well networked with, and liaises with Information Security to establish security classification of data
May require management skills to manage more junior Data Stewards
A Metadata Registry is a registry or repository used to not only store your data element definitions, but how those data elements relate to all other entities in your enterprise. The Metadata Registry need not be a sophisticated enterprise application (as many Metadata Registries will support ISO 11179 or OMG MOF data definitions, layered data models, Change Management user workflow, modelled business entities, integrations with ETL tools, BI tools, and IDEs, and more), but can be as simple as a spreadsheet or even text document. As long as the document has enough structure to show what each data element’s fully qualified definition is, but detailed enough to distinguish between similar but different data elments. In my experience the vast majority of Metadata Registries are just simple Excel spreadsheets, or at the most a Microsoft Access DB.
Alternatively, if you already have another type of repository such as a Configuration Management Database (CMDB), it may be possible to extend this registry to include data element definitions. By doing so, you extend the value of your original repository while at the same time solving the problem of where to store your data element definitions. From what I can tell, this is one of the best strategies you can follow as you are only building what you need while extending your IT value chain. By the way I shouldn’t take credit for this, this is really just my interpretation of a methodology and approach that has been proposed by Charles T. Betz and which is described in his book "Architecture and Patterns for IT Service Management, Resource Planning, and Governance: Making Shoes for the Cobbler's Children". You can also read his blog at: www.erp4it.com.
All right, so I’ve described what a Subject Area is, what Data Stewards are responsible for, and what a Metadata Registry is (you can also read more about Metadata in my earlier blog post: “Metadata Defined”). So, how does this all pertain to Data Governance? And how do I implement Data Governance?
When talking about any form of implementation, we are talking about the HOW and not the WHAT (I covered that in my last post). As we all know, there’s a million ways to skin a cat, so anyone who claims to know “the way” is fooling themselves. So, the steps described below, are merely “a way”, but a way that I have seen work in other companies with great success. To be perfectly honest, the steps I’m outlining below are a composite of many ways that I have observed, and I’ve just tried to find the common themes. So, in simple bullet form below, I am describing in ultra-simplistic detail HOW to implement data governance:
Determine which Subject Areas you would like to implement data governance for. It is recommended that you choose subject areas that have the greatest number of data contention issues, and therefore the greatest ROI for your governance of those Subject Areas. As this will provide you with the greatest opportunity for funding. Securing funding of course is often the hardest part in any new venture. However, given that you will be working with the most contentious data it’s difficult to wade into this slowly. Alternatively, if you can get funding for a pilot project that covers less contentious data then this might be the way to go from a risk perspective.
Put together a business case for project funding. This will involve determining the ROI for the initiative. For an outside consultant, this is difficult as you are not aware of the specific problems felt by lacking Data Governance. What you can attempt is a Data Governance Audit. As a matter of fact, IBM has recently announced a Data Governance Service. The IBM service is essentially an 11 point audit, I suspect followed up with appropriate recommendations. Selling an audit is often tough to do, but if you are going to pitch one, the best time to do so is at the end of the budget year when departments are looking for ways of spending their remaining cash (to ensure the same or greater budget for next year). Additionally, there are many factoids being published on a regular basis which you can use to justify Data Governance in general. For example, Accenture recently completed a study showing that middle managers waste two hours each day just searching for the right information, and once they get it, nearly half of it is useless.
Assuming you get funding, the next step is to determine the Data Owners of the respective Subject Areas. Determining who owns what data is probably the hardest part as there will be people who don’t want to take ownership of data, and others who want to take more ownership than they should be entitled to, and then of course there are those messy grey area situations where there may need to be multiple owners of data. The best way of convincing someone that they should own a Subject Area is to ask them how much they have to lose if the data quality is poor, or if someone else who doesn’t have as vested an interest in the data instead takes ownership. Yes, it’s a negative way of thinking but sometimes you need to paint a picture of anarchy and chaos to mobilize people into owning their data. This is especially why it’s important for you [the Data Management professional] to have excellent negotiation and persuasion skills.
After identifying and establishing Data Owners, it is now time to round-up the Data Stewards who will work on the Data Owners’ behalf to ensure that the Data Element definitions are correct. It is generally a good idea to select people who have a good knowledge of the data to begin with (as this will save time in sourcing and documenting Data Definitions), but also someone who has strong analytical skills. This could be a business analyst, a software developer, a DBA, a project manager, or potentially even a Customer Service Rep. Since the Data Steward role is just a role, it need not be a full time job. In fact most Data Stewards spend less than half of their time on Data Stewardship.
While selecting your Data Stewards you will also need to consider the reporting structure, and depending on how far you want to go, you may want to consider multiple tiers of Data Stewards. I have witnessed organizations with three levels of Data Stewards, but they had over 150 Data Stewards which is something that clearly does not happen over night.
You now have clearly established in scope Subject Areas with corresponding Data Owners. You also have Data Stewards assigned to work on behalf of those Data Owners and who are Subject Matter Experts (SMEs) for those Subject Areas. Now you need to select your Metadata Registstry. Software selection in general is never a trivial task, so I’m not going to pretend that selecting a Metadata Registry is any easier. There are best practices particular to selecting a Metadata Registry which I’ll write a later blog entry about, but for now will leave this one open, and just assume that it will happen. One thing I can reveal, is that if you’re just getting started with Data Governance then you’re best to keep the Metadata Registry simple and easy to use. Something like an MS Access DB or even an Excel spreadsheet will do. At a later point you can always migrate to a more robust solution. Alternatively, as I mentioned earlier, you can think about extending an existing repository such as a CMDB. However, if you choose this route, it will likely be a lot more difficult to migrate to a different solution, or change what you have.
Your ducks are all lined up now! You just need to tie it all together with policy and procedure. This is where you need to put your Soft Systems thinking hat on and figure out the best way of ensuring people can do their job building systems, repairing systems, and decommissioning them without too much policy and procedure getting in their way. But at the same time ensuring that the Data Owners’ best interests are being preserved. For this you will need to determine the following:
Which technology domains are in scope and which are out of scope. For example you don’t want to waste your time governing how people modify a tiny workgroup MS Access DB. Nor do you necessarily want to govern how a hash file is maintained. My advice is to concentrate first on data that is most likely to be shared across technology domains and business units. In particular I recommend focussing on RDBMS data elements (e.g. data contained within an Oracle, DB2 or MS SQLServer DB). To this day, the relational model is still the only complete data model. No other data model provides built-in guarantees of: referential integrity; value constraints; and access and update times. Even with XML, a document can point to any other document, even if that document doesn’t exist. Additionally, the RDBMS has more adaptors than any other data store, so more people have the ability to easily connect to it. Finally, there is already a great deal of rigour around the RDBMS, from back-up regimes, to security regimes, to access regimes. So data stored within the RDBMS is already perceived as more of a hardened asset than data stored in other forms.
You will need to determine what the new procedures will be for:
i. Adding new Data Elements
ii. Changing existing Data Elements
iii. Removing or decommissioning Data Elements
How this procedure will work in a multi-tier environment. I.e. how the process would work when there is a Development environment, a Test environment and a Production environment (or how many tiers you may happen to have).
How to integrate the procedure with exiting Change Management procedures and processes to ensure the correct approvals take place, but also ensuring that no more people are required to make a change than necessary
Who will be involved when executing the aforementioned procedures. Creating a RACI chart is a good way of documenting this.
If you are working for a smaller organization, then the above methodology can be greatly simplified. You will be able to quickly and efficiently determine who the Data Owners and Data Stewards are. Coming to an agreement on how the procedures will work, will also be a lot faster. As a smaller organization you would also be wise to keep your Metadata Registry very simple, such as an Excel spreadsheet or MS Access DB.
If you are working for a large organization, none of what I described above will come easy, but don’t look for silver bullets. Yes, good software can perhaps automate some of the manual procedures through people workflows and tool integrations. Good software may even help in the initial discovery process to document your data flows (especially when ETL tools are involved). However, most of the hard work comes back to the right people being properly engaged to make well informed decisions.
Going back to my first example where I talked about a lawyer working with you to complete a Real Estate transaction. Yes, you and the lawyer may be able to use technology to work more efficiently. For example, in the future (or maybe the present by the time you read this) your lawyer will be able to submit your deed to the Land Registry on-line, saving some time. Or maybe you and your lawyer can collaborate on-line without you having to make the trek down to her office to sign documents. Regardless, I would still say that the most critical work that is being done is you making sound decisions, and allowing your lawyer to interpret these decisions while ensuring that all the legal details are being taken care of in your best interests. Until artificial intelligence can make some great strides (which I’m not seeing), we’re going to depend on lawyers to help us make decisions when the law is concerned, and Data Stewards when data is concerned.
Posted by
Neil Hepburn
at
10:22 PM
0
comments
Monday, December 18, 2006
Rationale for Data Governance
When you think of Canada’s infrastructure you probably think of physical things like: highways; railways; copper, fibre optic, and coaxial communications lines; gas lines, electrical lines, and power generation; water filtration, and sewage treatment; solid waste management; airports; water ports; post offices; hospitals; schools; and so on. These things are all definitely part of our nation’s infrastructure and constitute a major part of the bedrock that allows people and businesses to build upon. Take away these things, and you’re forced to reinvent them on your own. Looking at the current situations in Afghanistan and Iraq is a constant reminder of how critical infrastructure is for a nation’s stability and prosperity, and how hard life is without them.
In addition to the infrastructures I’ve mentioned above there are also other less tangible national infrastructures that are arguably even more important: constitutional; legal; political; monetary policy and banking; social services & welfare; crime prevention, safety, & policing; and so on. Most people are also aware of these infrastructures (although they may not refer to them as infrastructure), and understand their intrinsic value without question.
Yet there is another piece of our infrastructure that I haven’t mentioned. This piece of our infrastructure allows our government to make informed fact-based decisions at all levels of government. Likewise, this piece of our national infrastructure also allows both local and international businesses to make informed decisions in the way of marketing and execution. What I’m talking about here is our Census and other national Surveys collected and managed by Statistics Canada.
“The total estimated cost of the 2006 census is $567 million spread over 7 years, employing more than 25,000 full and part-time census workers” (source: Wikipedia). Most people’s familiarity with the Census is through newspaper tidbits or soundbites like “Did you know that our population is blah blah blah now” or “Did you know that ethnic minorities constitue blah blah blah percent of our population now”. So for the average person the Census may seem like a giant project to fuel watercooler talk. In fact, not only does the government and most large business rely on the Census to make fact-based decisions, but there is an entire industry built around interpretting and deriving new facts from current and past Census data. Take for example Environics Analytics who apply the theory of “Geodemography” to derive Market Segments that allow businesses to target segments from the thrifty “Single City Renters” to the more affluent “White Picket Fences” to the ultimate “Cosmopolitan Elite”. Another company, CAP Index, uses Census data to predict crime levels for homicide, burglary, auto-theft, assault, and other forms of crime through the application of the theory of “Social Disorganization”, based on an interpretation of Census data. There are also companies like MapInfo that will extrapolate Census data to predict what Census data may look like today, or tomorrow. Another company, Global Insight, will summarize or “flatten” Census data for comparison against other countries.
I believe that fact-based decision making is beginning to captivate the popular imagination. Books like Moneyball which showed how the Oakland A’s went to the top of the MLB simply by making fewer emotional decisions, and shifting to a fact-based data driven approach are arguments that even the non-sports buff can understand and appreciate. Another book “Good to Great”, takes a fact-based approach to dispel myths about the “celebrity CEO” or the “cult of IT” as reasons for sustained corporate success. On the flip-side, Malcolm Gladwell’s Blink may seem to contradict this thinking in that it exemplifies the “intuitive expert”. However, the picture of “intuitive expert” that Gladwell paints is not just anyone. This person has an historical verifiable track-record in making decisions pertaining to a [typically narrow] domain. Thus, if we can find an expert who has a proven track record in making well-defined decisions (e.g. identifying forged art, or identifying marriages in decline), it is fair to let that person make intuitive decisions in this domain based on this person’s fact-based track record. If such an expert cannot be found, we must search elsewhere for our facts.
Intuition is essential in business, but the organization who knows where and when to apply intuition and justify the application of intuition through facts will have a better chance of success than the organization that does so in an unchecked manner. In other words, don’t get your “corporate heroes” that have had a generally successful track-record mixed up with your “intuitive experts”. I’ve gone off on a bit of a tangent here, but the point I’m trying to make is that high quality data is essential to both the government and free enterprise for making strategic decisions.
Okay, so the Census is important. So what?!? Now, I can get to the point I really want to make: The Census is only possible through strict data governance.
Not surprisingly, most people can’t be bothered relinquishing private information, nor can they be bothered filling out forms that take over an hour to complete. So, back in 1918 the Canadian government passed the Statistics Act which as it stands today, makes NOT filling out the Census, or providing false information punishable by up to a $1000 fine and 6 months in prison (source: Statistics Canada). With such legislation in place, it is in ones own best interests to complete the Census.
On the other side of the Census coin, if we look at Census data for Aboriginal communities (aka First Nations communities). By Statistics Canada’s own admission, there is a dearth of quality Census data. In fact, this lack of quality data has been cited specifically as a major obstacle to economic improvement in First Nations communities. As a consequence, the government of Canada is actively addressing this issue through the First Nations Fiscal and Statistical Management Act (FSMA) which was introduced in the house of commons December 2nd 2002. The FSMA calls for a dedicated body referred to as the First Nations Statistical Institute (aka First Nations Statistics [FSN]). This body will be integrated with Statistics Canada, and over time I hope to see some improvement with the quality of Census data coming out of First Nations Reserves. As matter of fact, I was browsing through the business section of The Globe and Mail a few weeks ago and noticed four advertised positions for statisticians to work in the FSN, so clearly things are moving forward, although probably slowly given it’s a government initiative.
I’ve met some [cynical] people who think that this improved data will make no difference to the lives of those on the reserve. While I can’t say that the improved data will make First Nations lives better, I can say with certainty that the lack of quality data is an obstacle to improvement. Furthermore, the Census can highlight communities that are working which could serve as reference models. The Census data can also definitively show what communities need the most assistance. The data can free up many political logjams, as the conversation is allowed to move from highly politicized discussions of the towns themselves (which are typically ego driven) to more rational discussions about what Critical Success Factors (CSFs) or Key Performance Indicators (KPIs) to look for, and what the definition of those CSFs and KPIs are. Since the data itself is currently dubious, those discussions quickly get derailed. But when the data does become reliable, we’ll have the bedrock for such a dialog.
The corporate world is no different. The vast majority of businesses have poor or non-existent data governance. When arguments flare up over the meaning of data, someone is quick to point out flaws in the quality of data itself, and will use this as a “hostage for fortune” to push their own agenda.
Okay, so I’ve talked about the [Canadian] government, and briefly mentioned corporation’s similar woes, but what about the wild-west that is the internet?
Wikipedia, as many of you know by now is the worlds largest “open source” encyclopaedia. Wikipedia is one of the best examples of the new emerging collaborative culture that’s sweeping the internet (aka Web 2.0). I personally love Wikipedia. While I am not an expert in everything, for those things that I do feel I know more than the average person about, I’m astounded by the amount of detail, and even accuracy of information on Wikipedia. However, I am also aware of its flaws. Some of the flaws are obvious, but believe it or not, the biggest issue with Wikipedia is rarely mentioned.
The most highly publicised flaw Wikipedia has been fingered for is maintaining the so-called “Neutral Point of View” or NPOV as Wikipedia calls it for short. Highly politicized subjects tend to create a “tug-of-war” of facts and opinion. At the time of this writing, the list contained 4,994 English articles lacking neutrality out of a total of 1,528,201 articles. So, approximately 0.3% of all articles are “disputed”. It’s actually quite interesting to see what they are. Topics range from the usual suspects: “Abortion”, “Israeli Palestinian violence”, to the more staid “SQL”, or “Audi S3”. You can read about the NPOV here:
http://en.wikipedia.org/wiki/Wikipedia:NPOV_dispute
You can also get a list of all articles with questionable neutrality here:
http://en.wikipedia.org/wiki/Special:Whatlinkshere/Wikipedia:NPOV_dispute
However, to get a better idea of the actual likelihood of someone clicking on an NPOV article, I took a look at Wikipedia’s own statistics. Namely, the 100 most requested pages. Of the 100 most requested pages for the month of December 2006, 88 of these pages are articles (the other 12 are utility pages, like the search page). I took a look at each article and noticed that no fewer than 14 are “padlocked” meaning that they can only be edited by veteran Wikipedia users – no newbies allowed. So, of the 88 most requested articles for Dec. 2006, 16% are strictly governed. I suspect that all of these articles at one time had NPOV issues, since padlocking an article is the best way of stopping the tug-of-war.
So the Neutral Point of View is surely a serious issue, but because Wikipedia has flagged these articles, there is a cue for the reader to approach them with a more critical mindset.
Perhaps a bigger problem is with Wikipedia vandals. People who make phoney edits just for the heck of it. Stephen Colbert famously encouraged his loyal viewers to do this, which they did. This not only resulted in Colbert being banned from Wikipedia, but also a lock-down on his own page, ensuring that only vetted users could modify its contents. Furthermore, Colbert’s stunt was cited as a major reason for starting Citizenddium .
Citizenddium was started by Wikipedia co-founder Larry Sanger as more strictly governed version of Wikipedia. I’ve quoted the following line from this CNET article on Citizendium, which I think is very telling of some of the issues faced by Wikipedia:
But unlike Wikipedia, Citizendium will have established volunteer editors and "constables," or administrators who enforce community rules. In essence, the service will observe a separation of church and state, with a representative republic, Sanger said.
Looks like data governance is only getting stronger here. So much for the wild-west. I can’t say I’m really celebrating this because I realize that Wikipedia wouldn’t be were it is if it started off as a rigid members only club.
However, I haven’t yet mentioned the biggest problem that Wikipedia faces. Namely: vocabulary. You can read more about this issue here: http://meta.wikimedia.org/wiki/Governance
For your convenience, I’ve quoted the most pertinent paragraph:
A common concern is that of our vocabulary, which necessarily expands to deal with professional jargons, but must be readable to more casual users or those to whom English is a second (or less familiar) language. This affects directly who can use, or contribute to, the wikipedia. It's extremely basic. Unless it's dealt with, you aren't able to read this material at all.
Interestingly enough, this particular issue maps almost directly back to the issue of Metadata (or lack thereof) that most corporations are still grappling with (see my previous post on Metadata). Therefore, once again, on this matter if Wikipedia is to properly address this issue, it must do so through improved governance. Sure software will hopefully create workflows to ensure that policies are being executed efficiently, but to be sure this is not an issue that can be solved by software alone.
Clearly, the quality of data hosted by Wikipedia is not at the “bet your business” level for all articles, and that in order for it to get there, more governance, or a complete governance overhaul (such as what Citizendium is doing) is required. In spite of these issues, I would still categorize Wikipedia as huge success in its own rite, and a model that may appear very attractive to a maverick in the corporate world.
However, I should point out one major but non-obvious difference between Wikipedia’s data and corporate data. Wikipedia’s data is essentially “owned” by the same people who go in and physical modify the articles. In other words, The Data Owner is the Data Steward. People are only adding or changing Wikipedia data because they themselves have a vested interest in those data. It is therefore in the Data Owner’s best interest to ensure the data being entered into Wikipedia is as accurate as possible. Furthermore, because all changes are done so on a volunteer basis out of self-interest, there is no “Triple Contraint”. Namely, there are no trade-offs that are required between: cost; quality; and time. Thus, the Data Owner can have [according to her own desires] the highest level of quality maintained at all times. Otherwise, there’s no point in making an edit to the article.
Enterprise data does not share the same luxury as the Wikipedia model. Those that are managing the data are usually in IT and are not the same as the Owners of the data who typically reside in other business units. Therefore, it is important to ensure that whoever makes changes to, or utilizes data, follow strict guidelines to ensure that the Data Owner’s interests are being met, and that quality be maintained. Left to ones own motivations, people will take the path of least resistance which over time will lead to degradations in Data Quality, and Data Interoperability, as the “Triple Contstraint” will force these issues. Taken further out, this will lead to increased costs. Both: tangible (e.g. increase in trouble ticket resolution times); and intangible (irate customers).
To sum up, we know that good governance over our data, not only makes it an asset, but can even be thought of as an investment. The Canadian Census did not happen by accident, but through a rigorous governance model, with real world penalties such as fines and jail time. As a result, Canada enjoys many economic advantages over countries that do not have the same quality of data. On the other hand, it is possible to have acceptable data quality in weakly governed environments, but those environments really only thrive if they are controlled by their Data Owners, and those owners are not bound by the “Triple Constraint”. However, even in the most utopian of environments, there must still be some level of governance to ensure highly consistent levels Data Quality and Data Interoperability (shared and understood vocabulary). If you want high levels of Data Quality and Data Interoperability, you cannot do so without creating policies, assigning roles, and implementing procedures. There are no Silver Bullets, and Data Governance is a necessity.
Posted by
Neil Hepburn
at
12:19 PM
0
comments
Labels: data governance metadata
