ASUG News + Views
Deprived Mas­ter Data and Its Remedies
Bala Subbaih Oct 5, 2019
Bookmark
Share Article:

Envi­sion a com­pa­ny that has bad data but per­fect process map­ping. A com­pa­ny with this set­up could lose rev­enue because inac­cu­rate data may lead to chal­lenges such as high­er resource con­sump­tion, high­er main­te­nance costs, neg­a­tive pub­lic­i­ty on social media, and low­er pro­duc­tiv­i­ty. The com­pa­ny would need to erase its old data and col­lect new data, which like­ly would require spend­ing more time and mon­ey. This is a red flag.

On the oth­er hand, a com­pa­ny with good data but bad process map­ping may still need to spend time attempt­ing to rec­ti­fy these process­es. This com­pa­ny may need to real­lo­cate resources to improve data qual­i­ty, yet this sce­nario will not be as expen­sive as fix­ing down­right bad data.

Accord­ing to Gart­ner research, orga­ni­za­tions believe poor data qual­i­ty to be respon­si­ble for an aver­age of $15 mil­lion per year in loss­es.” Larg­er orga­ni­za­tions that work with more cus­tomer, employ­ee, sup­pli­er, and prod­uct data are at a high­er risk of encoun­ter­ing poor data quality.

Sev­en Key Data Concepts

As a mas­ter data prac­ti­tion­er, I have some insight in how to work with data that I’d like to share through these sev­en key data concepts.

Fat Records: Busi­ness­es col­lect a lot of attrib­ut­es based on their require­ments. Depend­ing on the busi­ness, the mate­r­i­al mas­ter, arti­cle mas­ter, busi­ness part­ner, and all oth­er assets need dif­fer­ent attrib­ut­es. For exam­ple, mate­r­i­al mas­ter data typ­i­cal­ly uses more than 250 attrib­ut­es, which may include details such as plant data, sales data, pur­chas­ing data, account­ing data, ware­house data, etc.

A fat record will hold many attrib­ut­es. If you are unable to man­age your busi­ness-crit­i­cal attrib­ut­es effec­tive­ly due to their quan­ti­ty and unstruc­tured nature, the busi­ness and data cus­to­di­an might not know the sig­nif­i­cance of the busi­ness-depen­dent attrib­ut­es. It is extreme­ly impor­tant that the attrib­ut­es of any busi­ness object are kept concise.

Fed­er­a­tion of Ele­ments and Records: Hav­ing an orga­nized view of ele­ments and fed­er­a­tion of func­tion­al ele­ments are key con­cerns for data cus­to­di­ans and the busi­ness as a whole. A busi­ness entity’s group­ing, clas­si­fi­ca­tion, and hier­ar­chy can make assess­ment, audit­ing, and over­haul­ing sim­ple. Cen­tral data ele­ments of the object should be at the top of the schema and then cat­e­go­rized by func­tion­al grouping.

Half-Fin­ished Data: Infor­ma­tion should be com­plete. Miss­ing data can lead to mis­lead­ing analy­sis and results. Busi­ness­es must spend time and resources as they attempt to recre­ate and recov­er lost data. Some busi­ness­es may not be will­ing to do so and, instead, may leave these records unused. This leads to set­backs in pro­duc­tiv­i­ty time­lines, loss of trust from cus­tomers, and data repair/​replacement costs.

Data is a dig­i­tal asset. Every part of this infor­ma­tion is valid and indis­pens­able. Sup­pose we have incom­plete equip­ment mas­ter data, where the core func­tion­al attrib­ut­es such as type, func­tion, capac­i­ty, and age are left unde­fined. This leads to ambi­gu­i­ty in the use of the equip­ment and could harm the business.

Own­er­ship and Account­abil­i­ty: Unsteady own­er­ship and account­abil­i­ty lead to half-fin­ished and unused records. Busi­ness­es with cus­tomized, effec­tive gov­er­nance mod­els will have high­er-qual­i­ty data because man­age­ment respon­si­bil­i­ties, roles, account­abil­i­ties, data flow, and oth­er guide­lines are strict­ly defined and put into action. Gov­er­nance oper­at­ing mod­els improve data coor­di­na­tion, leav­ing lit­tle room for mistakes.

It is high­ly encour­aged to improve data qual­i­ty by estab­lish­ing struc­tured schemas, a hier­ar­chy of respon­si­bil­i­ties, process flow doc­u­men­ta­tion, and a clear set of to do and not to do” tips.

Track and Pro­file: To avoid incon­sis­tent, inap­pro­pri­ate, or half-fin­ished data, we need data pro­fil­ing and change track­ing. This will help you keep track of the who, what, when, and why that’s need­ed to main­tain qual­i­ty data.

Meta­da­ta, Schema, and Mod­el: The fore­most objec­tive of a schema data mod­el is to main­tain an accu­rate, com­pre­hen­sive rep­re­sen­ta­tion of the objects in the appli­ca­tion. A poor data mod­el leads to deprived data. Every data object has its own schema attrib­ut­es and set of keys. Objects need to be keyed accu­rate­ly — whether the keys are pri­ma­ry or for­eign. We then need to iden­ti­fy and define the rela­tion­ships and asso­ci­a­tions between dif­fer­ent objects in the entire data­base, cre­at­ing orga­nized schemas.

In an enter­prise, objects can be a mul­ti­tude of things. So, it is cru­cial that detailed infor­ma­tion regard­ing each enti­ty is stored in the data­base and char­ac­ter­ized into var­i­ous fields, called attrib­ut­es, with details that are then organized.

Meta­da­ta is data about data, such as char­ac­ters, texts, or numer­als. Once again, to avoid incon­sis­ten­cy, the mod­el should have ref­er­ence val­ues, key map­ping, hier­ar­chy, clas­si­fi­ca­tion, group­ing, etc.

Dupli­cate Records: There are a lot of sce­nar­ios through which records in a sys­tem can be dupli­cat­ed. Dupli­cate records and val­i­dat­ing dupli­cates also can result in wast­ed time. Dupli­cate records may result from unpre­dictable source data and inad­e­quate ref­er­ence and hier­ar­chy data, due to het­ero­ge­neous sys­tems such as the var­ied char­ac­ter­i­za­tion of spe­cial char­ac­ters, punc­tu­a­tion, noise words, abbre­vi­a­tions, and some iden­ti­fi­ca­tions. Trans­form­ing and sub­sti­tut­ing such sources/​buzzwords may reduce the risk of cre­at­ing duplicates.

Sev­en Key Sources of Dupli­cate Occur­rences Dur­ing Runtime

  • Lack of own­er­ship and accountability
  • Lack of skills (own­er­ship comes from skills)
  • On-flight urgency
  • Incon­stant change track­ing and monitoring
  • Absence of data profiling
  • Bad configuration/​setup
  • Not hav­ing real-time data enrich­ment through third-par­ty systems

Defin­ing a desired lev­el of match­ing across records and iden­ti­fy­ing dupli­cates can help cor­rect and avoid dupli­cate entries.

Stay­ing Focused on High-Qual­i­ty Data

In the dig­i­tal age, com­pa­nies seem to be cold-shoul­der­ing the qual­i­ty of resources — rel­e­vant and skilled — as well as eth­i­cal and tra­di­tion­al tech­niques. They’re sim­ply look­ing at new tools and tech­nolo­gies to acquire data. Yet the qual­i­ty of its infor­ma­tion is a key to suc­cess for any orga­ni­za­tion. Avoid­ing these caus­es of poor data and bad process­es can help us all focus on improv­ing our data quality.

Reg­is­ter for the ASUG Expe­ri­ence for Enter­prise Infor­ma­tion Man­age­ment (EIM) Oct. 28 – 30 in Min­neapo­lis to learn from peers how to man­age your data to meet busi­ness goals. Addi­tion­al­ly, we wel­come all ASUG mem­bers to sub­mit their ideas for blog posts they want to write. 

You Might Be Interested In


Insights Included in Membership
View All Insights
Bookmark
Bookmark
Bookmark
Bookmark