Types of Master reference data
I have been pulled into some detailed discussions on Master reference data management and through this it occurred to me that there are different types of data that need a different treatment. Here are a few examples:
- The first type I would like to call 'external standard lists' - this is the reference data for which we already have (international) standards (like ISO). Examples are Countries, Post codes, Units of Measure, Language codes. Characteristics: Standard is defined, source is outside the company, no debate necessary. Approach: Just identify the data type and agree to use the standard from now on. Even better is to make it centrally available as a service.
- The second type is similar - but I would call them the 'internal standard lists' - things inside the company for which you like to have an agreed enumerated list - i.e. a fixed domain. Think about standard lists for security levels, status codes, ... Characteristics are: Standard is usually a lot harder to define due to debate from different stakeholders, while this debate is not really adding a lot of value. Approach: Identify the data type, propose a standard after analysis. Challenge the parties to make a case if they don't think they can comply to this standard. Make the standard lists available as a service.
- Third is a type you should skip and this is the set of items that people think as common attributes, but which are actually not manageable. Think about document types (everybody has a different opinion), keywords, names of project teams (too many!), etc. It is hard to define this category, but it is usually about lists that change constantly, that do not have a clear source or owner and where there are no clear definitions (or where it is just over the top to manage it as reference data!). I see a folksonomy as a good alternative for this category.
- Finally the fourth is the most important category - and that is for objects / things / people shared across the company the 'master reference objects' - this data you probably already maintain in one place or another and the trick here is to get agreement where you have the master etc. (stuff I covered before). Note that these objects can be organised in hierarchies - which in themselves can be seen as reference data (but lets leave that for now!)
Labels: folksonomy, master reference data
Master reference data and Folksonomy
Building up master reference data in terms of key objects in the company is difficult, but doable. Building up a reference data list in terms of 'information types' (e.g. document types, or other taxonomy) is even more difficult. From experience I can say it is usually a disaster (too much detail, too complex ...). Therefore the concept of a managed
Folksonomy is very compelling alternative. It is a Darwinian answer to data management (let the fittest document types survive). The approach could be as follows:
- Data managers can start with a first cut list of information types and other attributes available for users to pick and choose - the data managers can even do a bit of a mapping exercise to some of the documents, so at least some information is linked to the information types
- Then users are invited to use the document types for their own publication processes (probably with an accellerated speed if a document type is a mandatory attribute). Allow users to choose from a picklist or to create their own if they don't know what to choose
- Create statistics and remove the taxonomy entries with a low number of links - invite the users to reclassify if they have used the removed items
Through this gradually the list of types will grow and they will become fit for purpose.
Labels: folksonomy, master reference data
Taxonomy & Tagging
People managing information have been trying to find the ultimate taxonomy for many many years. Various mostly hierarchical structures exist that help us finding our way through libraries, collections of equipment, pharmaceuticals, etc. This is very helpful indeed when managing large collections of information, but usually the creative users considers these taxonomies as a nightmare. And if they can avoid it, they will not do it.
I believe (as stated in the previous post) that Search helps us mastering a lot of the information management problems we have today (around finding the information), but I also believe that we need to give Search a helping hand. So how do we do this?
The first thing a corporate tiger can think of is adding more
control - i.e. forcing people to use taxonomies, but this will not help a lot. You can of course fix top levels of directory structures and through this enforce some sort of taxonomy, but too much control will frustrate people and have them looking for loop holes.
Therefore we have to trust the user. We have to start relying on their own sense of responsibility. They will tag their documents in their own way. It is not only about trusting the users that they will do it (and in a 'good' way). It is also about trusting that this relative anarchy will give some order at the end of the day.
- About trusting the users: Usually people are not brainless. With other words they will try to do something sensible. Especially if they see that it is easy and that it helps others.
- About trusting it will work: Just look at the whole Web2.0 phenomena, as mentioned a few days ago. We have lots of self-organising sites already today and they work. Why not do this inside companies as well?
You also need to
help the users a little, just like in this Blog. If I want to use tags I can see what I have used before. At a company scale it would be even possible to suggest common tags to users (almost like taxonomy), but these tags do not exist because they were invented in an invory tower, but because they were used on a regular basis by others. This is called Folksonomy.
Labels: folksonomy, search, tagging, taxonomy, Web2.0
IM and the world of Web 2.0
I just watched a video on Web 2.0 (just search for Web 2.0 on
YouTube or download via
http://www.mediafire.com/?6duzg3zioyd) and I realised that most companies still live in the world of yesterday. Most information is hidden away in personal drives and protection is the keyword. Only a few people realize that most of the potential of all this hidden information is lost by the fact that it is overprotected and hidden.
I think companies can learn from the
Web 2.0 phenomena in rethinking the way they manage their information. This whole wave of new(ish) thinking can also be applied within the firewall of a company. Think about:
- Why do we have everything protected? My view is that everything should open up. Only a few parts of the infrastructure are only for named individuals (think contracts, think HR), but the rest should be open for everybody. This will unlock information in an easier way and will reduce the burden of managing security
- Why do we try to classify information with complex taxonomies or other classification schemes? People normally just fail to do this, so allow people to add their own tags and ensure these tags are easily visible. This will create a folksonomy for the company over time. Folksonomies are not perfect, but have much more potential than a top-down classification structure (just check out del.icio.us or flickr)
- Why don't we have quick company-wide search? Search is obviously one of the answers as well.
- Why don't we collaborate on content in a transparent way? Think about Blogging! (even the CIA is doing it these days!), but also think about creating a corporate memory via wiki's (http://www.wikipedia.org/).
Key answer for all questions is: Because most people managing information still live in the world of yesterday.
Labels: folksonomy, search, Tagging, Web2.0, wiki